Skip to content
#

serving

Here are 20 public repositories matching this topic...

Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.

  • Updated May 8, 2025
  • C++

End-to-end LLM serving simulator integrating scheduling, prefix caching, tensor allocation, and KV-cache management. 168-run sweep (72 baseline + 96 pressure). Key finding: ChunkedPrefill + LFU cache achieves 41% lower TTFT p95 and 94% prefix hit rate, but hits OOM first under memory pressure.

  • Updated Jul 10, 2026
  • C++

Simulates a request passing through 9 pipeline stages of an LLM serving stack (admission, queue, prefix cache, KV alloc, tensor alloc, prefill, KV transfer, decode, release), measuring where each millisecond goes. 15 configurations compared. Connects all 11 prior projects into a single instrumented pipeline.

  • Updated Jul 10, 2026
  • C++

Tracks each LLM serving request through the full pipeline with 24 event types, microsecond timestamps, and anomaly detection. Covers all state transitions: ARRIVAL -> ADMITTED -> QUEUED -> PREFILL -> DECODE -> COMPLETED, including prefix cache hit/miss, KV alloc, preemption, and KV transfer.

  • Updated Jul 10, 2026
  • C++

Simulates prefill/decode disaggregation for LLM serving across 192 configurations. Key findings: disaggregation beats monolithic in 19% of configs (requires arrival>=20 req/s AND prompt>=1024 tokens); best case 29% TTFT improvement; KV transfer scales linearly with prompt length; bandwidth has diminishing returns past 50 GB/s.

  • Updated Jul 10, 2026
  • C++

Add this topic to your repo

To associate your repository with the serving topic, visit your repo's landing page and select "manage topics."

Learn more