Skip to content
#

llm-serving

Here are 501 public repositories matching this topic...

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

  • Updated Oct 7, 2026
  • Python
ai-infra-handbooks

Eleven interlinked Chinese-language handbooks for LLM inference (AI Infra): Python, C++, CS fundamentals, math, LLM internals, CUDA, distributed training, inference systems, mini-sglang from scratch, image & video generation, and SGLang design evolution. Every example is auto-verified, and each chapter has exercises graded in the browser.

  • Updated Oct 7, 2026
  • Python

SGLang 在线推理服务三次挑战完整交付(学号 0102603133):HW1 部署与 Mooncake trace 压测、HW2 RadixAttention 前缀缓存测量与请求流程分析、HW3 四副本 Ray Serve 路由对比与改进(A/B/C/D)。含可复现脚本、逐请求结果与报告 PDF。

  • Updated Oct 7, 2026
  • Python
llmash

Ollama-compatible LLM server on a tuned llama.cpp fork. Same commands, API and model store as Ollama, up to 4x Ollama's decode speed and 2.8x vLLM's in our benchmarks, from custom CUDA kernels, MTP speculation and its own 3-bit builds. Uncensors models in one command. Windows and Linux, NVIDIA.

  • Updated Oct 7, 2026
  • C++

Add this topic to your repo

To associate your repository with the llm-serving topic, visit your repo's landing page and select "manage topics."

Learn more