kv-cache
Here are 1,096 public repositories matching this topic...
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
-
Updated
Oct 1, 2026 - Python
A Golang implemented Redis Server and Cluster. Go 语言实现的 Redis 服务器和分布式集群
-
Updated
Sep 14, 2025 - Go
A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
-
Updated
Sep 28, 2026
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
-
Updated
Sep 30, 2026 - Python
Unified KV Cache Compression Methods for Auto-Regressive Models
-
Updated
Aug 13, 2026 - Python
Official SGLang x Datawhale course on LLM inference: understand inference, build a mini-sglang from scratch, then read the real SGLang source and land your first PR. Available in English and Chinese.
-
Updated
Oct 2, 2026 - Python
LLM KV cache compression made easy
-
Updated
Oct 1, 2026 - Python
KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM
-
Updated
Oct 1, 2026 - C++
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
-
Updated
Aug 19, 2026 - Python
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
-
Updated
Oct 3, 2026 - Rust
[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
-
Updated
Sep 29, 2026 - Python
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation (NeurIPS 2025)
-
Updated
Sep 19, 2026 - Python
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
-
Updated
Feb 24, 2025 - Jupyter Notebook
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
-
Updated
Jun 5, 2026 - Python
[NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.
-
Updated
Aug 1, 2024 - Python
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
-
Updated
Jun 22, 2026 - Python
Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.
-
Updated
Jun 17, 2026
[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
-
Updated
Oct 1, 2026 - Python
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
-
Updated
Apr 26, 2026 - C
Add this topic to your repo
To associate your repository with the kv-cache topic, visit your repo's landing page and select "manage topics."