serving
Here are 201 public repositories matching this topic...
Simulates LLM serving admission control strategies (None, MaxConcurrency, TokenBudget, SLOAware, Predictive) across 135 configurations. Key findings: tight token budget (4096) achieves 10x better goodput than loose budget; predictive control is the only strategy maintaining SLO compliance at 100ms; SLOAware is fundamentally unstable under sustained
-
Updated
Jul 10, 2026 - Python
Load generator and performance analyzer for LLM inference serving. Measures TTFT, TPOT, and throughput against vLLM/TGI/Ollama or a built-in mock, and flags serving anti-patterns with tuning advice.
-
Updated
Jun 26, 2026 - Python
A flexible, high-performance serving system for machine learning models
-
Updated
Mar 31, 2018 - C++
-
Updated
Apr 12, 2026 - HTML
Production-ready LLM inference server optimized for AMD ROCm/MI300X. OpenAI-compatible API, PagedAttention, continuous batching, tensor parallelism, INT4/INT8 quantization.
-
Updated
Jun 3, 2026 - Python
TIDAL: Toolkit for Inference, Deployment, Adaptation, and Learning for efficient AI systems
-
Updated
Jun 30, 2026 - Python
Iteration-level continuous batching scheduler comparing FCFS, SJF, Fair, and Preemptive policies on real GPU inference. Validated across 5 seeds: SJF reduces short-job TTFT by 81.5% at the cost of 34% lower Jain fairness.
-
Updated
Jul 14, 2026 - Python
Event-driven benchmark of KV-cache migration policies for LLM serving scale-down, measuring session drop rate, infrastructure linger, and user-visible pause across workload types and concurrent session counts.
-
Updated
Jul 24, 2026 - Python
Professional Python project: deploying and serving machine learning models.
-
Updated
Aug 1, 2026 - Python
Profiling-driven SGLang optimization for quantized agentic LLM serving
-
Updated
Sep 18, 2026 - Python
An end-to-end Data Engineering project implementing a Medallion Architecture for real-time ride-sharing analytics.
-
Updated
Apr 4, 2026 - Python
Open research / proof of concept: pinned-expert MoE serving on Apple Silicon (MLX). Concurrency gains measured; >RAM streaming is a slow feasibility demo, not a product. AGPL-3.0.
-
Updated
Oct 6, 2026 - Python
This is base project of bentoml and machine learning model with poetry env
-
Updated
Mar 16, 2023
A single-binary HTTP load tester for **model-serving endpoints** (or any HTTP service). No dependencies beyond
-
Updated
Jul 3, 2026 - Go
Best Free AI Video Enhancement Tools 2026 - Ultra HD Upscale & Repair 🚀
-
Updated
May 1, 2026
Add this topic to your repo
To associate your repository with the serving topic, visit your repo's landing page and select "manage topics."