♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️
-
Updated
Oct 7, 2026 - Go
♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Agentic Runtime
Self-hosted multi-user LLM platform for the NVIDIA DGX Spark: OpenAI-compatible API (vLLM + LiteLLM), per-user keys & budgets, self-service portal with playground and an AI support assistant.
A high-throughput and memory-efficient inference and serving engine for LLMs
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
Eleven interlinked Chinese-language handbooks for LLM inference (AI Infra): Python, C++, CS fundamentals, math, LLM internals, CUDA, distributed training, inference systems, mini-sglang from scratch, image & video generation, and SGLang design evolution. Every example is auto-verified, and each chapter has exercises graded in the browser.
SGLang 在线推理服务三次挑战完整交付(学号 0102603133):HW1 部署与 Mooncake trace 压测、HW2 RadixAttention 前缀缓存测量与请求流程分析、HW3 四副本 Ray Serve 路由对比与改进(A/B/C/D)。含可复现脚本、逐请求结果与报告 PDF。
Community maintained hardware plugin for vLLM on Huawei Ascend
Abliterate refusal behaviors from LLMs with the most advanced open-source toolkit, evolving with every run.
Detect suspicious actions in video with computer vision models for security monitoring and behavior analysis
Translate human intent into Claude-ready prompts for cleaner, faster LLM coding tasks
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Characterizing CUDA Graph and eager decode execution regimes in SGLang with Nsight Systems and regime-aware service modeling.
AI platform reference architecture for LLM serving on Kubernetes, with reliability, observability, failure testing, and evidence-bound engineering claims.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Ollama-compatible LLM server on a tuned llama.cpp fork. Same commands, API and model store as Ollama, up to 4x Ollama's decode speed and 2.8x vLLM's in our benchmarks, from custom CUDA kernels, MTP speculation and its own 3-bit builds. Uncensors models in one command. Windows and Linux, NVIDIA.
Independent replication of CacheScout (agent-aware KV-cache eviction) as a vLLM 0.31 scheduler plugin, benchmarked on an 8 GB GPU with an honest scorecard.
LLM inference in Rust - Metal & CUDA
To associate your repository with the llm-serving topic, visit your repo's landing page and select "manage topics."