A flexible, high-performance serving system for machine learning models
-
Updated
Oct 6, 2026 - C++
A flexible, high-performance serving system for machine learning models
A scalable inference server for models optimized with OpenVINO™
A flexible, high-performance carrier for machine learning models(『飞桨』服务化部署框架)
A high-performance inference system for large language models, designed for production environments.
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
TensorFlow Serving ARM - A project for cross-compiling TensorFlow Serving targeting popular ARM cores
TensorFlow Serving based on encrypted model, protect model files from being stolen
pytorch during training, libtorch during serving via gRPC
SecretFlow-Serving is a serving system for privacy-preserving machine learning models.
A simple tensorflow C++ REST API server
tensorflow serving client using brpc
This project implements a common rest server which can serve tensorflow-serving & xgboost models.
End-to-end LLM serving simulator integrating scheduling, prefix caching, tensor allocation, and KV-cache management. 168-run sweep (72 baseline + 96 pressure). Key finding: ChunkedPrefill + LFU cache achieves 41% lower TTFT p95 and 94% prefix hit rate, but hits OOM first under memory pressure.
Simulates a request passing through 9 pipeline stages of an LLM serving stack (admission, queue, prefix cache, KV alloc, tensor alloc, prefill, KV transfer, decode, release), measuring where each millisecond goes. 15 configurations compared. Connects all 11 prior projects into a single instrumented pipeline.
A flexible, high-performance serving system for machine learning models
Tracks each LLM serving request through the full pipeline with 24 event types, microsecond timestamps, and anomaly detection. Covers all state transitions: ARRIVAL -> ADMITTED -> QUEUED -> PREFILL -> DECODE -> COMPLETED, including prefix cache hit/miss, KV alloc, preemption, and KV transfer.
A flexible, high-performance serving system for machine learning models
Simulates prefill/decode disaggregation for LLM serving across 192 configurations. Key findings: disaggregation beats monolithic in 19% of configs (requires arrival>=20 req/s AND prompt>=1024 tokens); best case 29% TTFT improvement; KV transfer scales linearly with prompt length; bandwidth has diminishing returns past 50 GB/s.
To associate your repository with the serving topic, visit your repo's landing page and select "manage topics."