Production-tested Kubernetes deployment recipes for LLM inference using NVIDIA Dynamo.
No recipe for your model and hardware combination? Point an AI coding agent at this repository and ask it to author or adapt one; the repo's agent skills guide the process.
Prerequisites: This guide assumes you have already installed the Dynamo Kubernetes Platform. If not, follow the Kubernetes Deployment Guide first.
These recipes compare Dynamo performance features with benchmark results, each including both baseline and optimized deployment configurations:
| Model | Framework | Configuration | GPUs | Features |
|---|---|---|---|---|
| Qwen3-32B | vLLM | Disagg + KV-Router | 16x H200 | Disaggregated Serving + KV-Aware Routing — benchmark comparison with real-world Mooncake traces |
| Qwen3-VL-30B-A3B-FP8 | vLLM | Agg + Embedding Cache | 1x GB200 | Multimodal Embedding Cache — benchmark comparison showing +16% throughput, -28% TTFT |
These recipes demonstrate aggregated or disaggregated serving:
GAIE Column: Indicates whether the recipe includes integration with the Gateway API Inference Extension (GAIE) — a Kubernetes SIG project that extends the Gateway API for AI inference workloads, providing load balancing, model routing, and request management.
| Model | Framework | Mode | GPUs | Deployment | Benchmark | Notes | GAIE |
|---|---|---|---|---|---|---|---|
| Qwen3-32B-FP8 | TensorRT-LLM | Aggregated | 2x H100/H200/A100 | ✅ | ✅ | FP8 quantization | ❌ |
| Qwen3-32B-FP8 | TensorRT-LLM | Disaggregated | 8x H100/H200/A100 | ✅ | ✅ | Prefill + Decode separation | ❌ |
| Qwen3-32B-FP8 | vLLM | Disagg (Single-Node) | 8x A100 | ✅ | ✅ | 2× TP2 prefill + 1× TP4 decode, NixlConnector KV transfer | ❌ |
| Qwen3-235B-A22B-FP8 | TensorRT-LLM | Aggregated (Hopper) | 16x H100/H200 | ✅ | ✅ | MoE model, TP4×EP4 | ❌ |
| Qwen3-235B-A22B-FP8 | TensorRT-LLM | Aggregated (Blackwell) | 16x B100/B200 | ✅ | ✅ | MoE model, TP4×EP4, DEEPGEMM backend | ❌ |
| Qwen3-235B-A22B-FP8 | TensorRT-LLM | Disaggregated (Hopper) | 16x H100/H200 | ✅ | ✅ | MoE model, Prefill + Decode | ❌ |
| Qwen3-235B-A22B-FP8 | TensorRT-LLM | Disaggregated (Blackwell) | 16x B100/B200 | ✅ | ✅ | MoE model, Prefill + Decode, DEEPGEMM backend | ❌ |
| Qwen3.8-2.4T-A95B-FP8 | vLLM, SGLang | Aggregated / Disaggregated | 16x GB300 / GB200 | ✅ | ❌ | Hybrid gated-delta-net + 512-expert MoE (262K ctx), FP8 weights + FP8 KV, TP16 over MNNVL, KV-aware routing + prefix caching, reasoning + tool calling | ❌ |
| Qwen3.5-122B-A10B-NVFP4 | vLLM | Aggregated | 2x B200 | ✅ | ✅ | Hybrid GDN+MoE, NVFP4 + FP8 KV, TP1 x replicas: 2, KV-aware routing; agentic profile |
❌ |
| Qwen3.5-122B-A10B-NVFP4 | vLLM | Disaggregated | 3x B200 | ✅ | ✅ | Hybrid GDN+MoE, NVFP4 + FP8 KV, 1P2D over NIXL, KV-aware routing; agentic profile | ❌ |
| Qwen3.6-35B-A3B | SGLang | Aggregated | 1x B200/GB200/H200 | ✅ | ✅ | Hybrid GDN+MoE, NVFP4 + FP8 KV on Blackwell / FP8 weights + BF16 KV on H200, MTP spec decode, TP1 | ❌ |
| GPT-OSS-120B | TensorRT-LLM | Aggregated | 4x GB200 | ✅ | ✅ | Blackwell only, WideEP | ❌ |
| GPT-OSS-120B | TensorRT-LLM | Disaggregated | 5x Blackwell (GB200/B200) | ✅ | ✅ | Prefill/Decode split | ❌ |
| GPT-OSS-120B | vLLM | Agg + Disagg | 8x B200 / 8x H200 | ✅ | ✅ | MXFP4 MoE + FP8 KV, 8x TP1 agg / decode-heavy single-node disagg (2P6D B200, 4P4D H200), EAGLE3 spec decode, KV-aware routing, harmony reasoning + tool calling; agentic profile | ❌ |
| Qwen3.5-122B-A10B-FP8 | vLLM | Aggregated | 4x H200 | ✅ | ✅ | Hybrid GDN+MoE, TP2 x replicas: 2, MTP spec decode, KV-aware routing; agentic profile |
❌ |
| Qwen3.5-122B-A10B-FP8 | vLLM | Disaggregated | 3x H200 | ✅ | ✅ | Hybrid GDN+MoE, 1P2D over NIXL, KV-aware routing, no MTP; agentic profile | ❌ |
| GLM-5-NVFP4 | SGLang | Disagg Prefill/Decode | 20x GB200 | ✅ | ✅ | NVFP4, EAGLE speculative decoding, TP16 decode + TP4 prefill, stable SGLang runtime image | ❌ |
| GLM-5.2 | SGLang | Aggregated + Disaggregated | 16x/20x B200 or 24x/16x H200 | ✅ | ✅ | B200 NVFP4 or H200 FP8 with FP8 KV, KV-aware routing, EAGLE, B200 HiCache CPU offload, agentic trace profile | ❌ |
| A.X-K2 | vLLM | Aggregated / Disaggregated | 8x / 12x B200; 32x H200 | ✅ | ✅ | Two TP4 aggregate workers or 2P1D TP4 over NIXL; prefill async disabled in 2P1D, EAGLE3 k=3 on both roles, NVFP4 + FP8 KV, sparse MLA, FlashInfer autotuning disabled, KV-aware routing. H200: FP8, TP8 with expert parallelism, KV-aware routing, EAGLE3 k=3, 262K context; c16–128 AIPerf sweep. | ❌ |
| DeepSeek-R1 | SGLang | Disagg WideEP | 16x H200 | ✅ | ❌ | TP=8, single-node. Use model-download-sglang.yaml |
❌ |
| DeepSeek-R1 | SGLang | Disagg WideEP | 32x H200 | ✅ | ❌ | TP=16, multi-node. Use model-download-sglang.yaml |
❌ |
| DeepSeek-R1 | TensorRT-LLM | Disagg WideEP (GB200) | 36x GB200 | ✅ | ✅ | Multi-node: 8 decode + 1 prefill nodes | ❌ |
| DeepSeek-R1 | vLLM | Disagg DEP16 | 32x H200 | ✅ | ❌ | Multi-node, data-expert parallel | ❌ |
| DeepSeek-V4-Flash | vLLM | Agg + Disagg | 4x B200 / 4x H200 | ✅ | ✅ | Text — MoE 284B / 13B active, NVFP4 (B200) / public FP8 (H200) + FP8 KV, agg TP4 (B200) / DP4+TP1+EP (H200), MTP (H200), KV-aware routing, agentic trace profile, reasoning + tool calling; plus disagg 2P1D (12x B200) / 4P3D (28x H200) | ❌ |
| DeepSeek-V4.1-Flash | SGLang | Agg + Disagg | 8x GB200 | ✅ | ❌ | Day-0, not benchmarked — MoE, FP8 dense + FP4 experts + FP8 KV, 1M ctx, TP4 + EP4, KV-aware routing, DSpark spec dec (agg only), reasoning + tool calling; plus disagg 1P1D with Mooncake KV transfer over MNNVL | ❌ |
| DeepSeek-V4-Pro | vLLM | Agg + Disagg | 8x B200 / 8x H200 | ✅ | ✅ | Text — MoE 1.6T / 49B active (1M ctx; 86k on H200), NVFP4 (B200) / public FP8 (H200) + FP8 KV, TP8 + EP, MTP-2 (B200), KV-aware routing, agentic trace profile, reasoning + tool calling; plus disagg 1P1D (16x B200) / 1P3D (32x H200) | ❌ |
| DeepSeek-V4-Pro-0813 | vLLM | Agg + Disagg | 8x GB200 / 8x H200 | ✅ | ✅ | Text — MoE 1.6T, MXFP4 experts + FP8 KV, 1M ctx, TP8 + EP, KV-aware routing, DSpark spec dec (GB200), agentic trace validated. Distinct checkpoint from DeepSeek-V4-Pro — do not share a model cache. | ❌ |
| Kimi-K2.5 | TensorRT-LLM | Disaggregated | 24x GB200 | ✅ | ✅ | DEP4 prefill + TEP4 decode, TRTLLM-native KV host offload | ❌ |
| Kimi-K3 | vLLM | Agg + Disagg | 16x GB200 / 16x GB300 | ✅ | ❌ | Multimodal MoE (1M ctx), MXFP4 experts + BF16 dense + FP8 KV, TP16 over MNNVL (GB200) / TP8 (GB300), KV-aware routing, FlashInfer MLA, reasoning + tool calling; plus disagg 1P1D (32x GB200) / 1P2D (24x GB300) | ❌ |
| Kimi-K2.6 | vLLM | Aggregated | 4x B200 / 8x H200 | ✅ | ✅ | MoE, NVFP4+FP8 KV (B200) / INT4 (H200), TP4/TP8, EAGLE3 MLA spec decode, LMCache CPU offload; text+image, chat + agentic profiles | ❌ |
| Nemotron-3-Super | vLLM | Aggregated | 4x B200 / 4x H200 | ✅ | ✅ | ~120B hybrid Mamba/Attention/MoE (~12B active), NVFP4 (B200) / FP8 (H200) + FP8 KV, TP4+EP, MTP, KV-aware routing; chat + agentic profiles | ❌ |
| Nemotron-3-Ultra | vLLM | Agg + Disagg | B200 / GB200 / H200 | ✅ | ✅ | Optimized agentic profiles for native 256K and opt-in 1M context; NVFP4 + FP8, MTP, and KV-aware routing | |
| Qwen3.8-Flash-Next | vLLM | Agg + Disagg | 4x B200 / 8x B200 / 12x B200 | ✅ | ✅ | Multimodal (text+image+video) ultra-sparse MoE (125B / 6B active) with GDN+QSA hybrid attention, 51B N-gram embedding offload to host RAM, Inferact NVFP4 weights, TP4+EP, MTP3 spec decode, KV-aware routing, reasoning + tool calling (qwen3_coder); agentic profile |
❌ |
| K-EXAONE 2.0 750B-A37B-NVFP4 | vLLM | Aggregated | 4x B200 | ✅ | ✅ | 764.5B MoE (37B active), hybrid attention (20 full + 58 sliding), NVFP4 W4A4 + FP8 KV, TP4, MTP spec decode, FLASHINFER_CUTLASS MoE backend required for correctness | ❌ |
| K-EXAONE 2.0 750B-A37B-NVFP4 | vLLM | Disaggregated | 8x B200 | ✅ | ✅ | Same model, 1P1D over NIXL/UCX on InfiniBand RDMA, MTP on both roles, decode --max-num-seqs 256 |
❌ |
Legend:
- Deployment: ✅ = Complete
deploy.yamlmanifest available - Benchmark: ✅ = Includes an AIPerf benchmark manifest
These recipes demonstrate functional deployments with Dynamo features, but have not yet been performance-tuned or paired with benchmark manifests.
| Model | Framework | Mode | GPUs | Deployment | Notes |
|---|---|---|---|---|---|
| Qwen3-32B | vLLM | Disagg 1P1D Provider Overlays | 8x A100/H100/H200/B200/GB200 | ✅ | Kustomize overlays for AWS EFA, GKE RoCE, AKS/Nebius/Nscale IB |
| Nemotron-3-Super-FP8 | vLLM | Aggregated | 4x H100/H200 | ✅ | TP=4, KV-aware routing |
| Nemotron-3-Super-FP8 | SGLang | Aggregated | 4x H100/H200 | ✅ | TP=4, KV-aware routing, 1.0+ |
| Nemotron-3-Super-FP8 | TensorRT-LLM | Disaggregated | 4x H100/H200 | ✅ | TP=2 prefill/decode split, UCX KV transfer |
| Nemotron-3-Super-FP8 | SGLang | Disaggregated | 4x H100/H200 | ✅ | TP=2 prefill/decode split, nixl KV transfer, 1.0+ |
These recipes are under active development and may require additional setup steps (e.g., container patching). They are functional but not yet fully validated for production use.
| Model | Framework | Mode | GPUs | Deployment | Notes |
|---|---|---|---|---|---|
| Gemma-4-31B | TensorRT-LLM | Aggregated | 8x B200 / 8x GB200 / 8x H200 | ✅ | NVFP4 + FP8 KV on B200/GB200; BF16 + 16-bit KV and TP4 on H200; KV-aware routing; MTP on B200/GB200, not yet verified on H200; agentic profile |
| GLM-5-NVFP4 (EFA) | SGLang | Disagg Prefill/Decode over AWS EFA | 20x GB200 | ✅ | KV transfer over AWS EFA via NIXL LIBFABRIC instead of UCX. Patched libfabric baked into image. Requires custom container build. |
| Nemotron-3-Nano-Omni-NVFP4 | vLLM | Aggregated | 1x GPU | ✅ | Multimodal text/image/video/audio serving. Requires custom container build. |
| nvidia/Kimi-K2.5-NVFP4 | TokenSpeed | Aggregated | 4x B200 | ✅ | Text only — MoE model, TP4×EP4, reasoning + tool calling. Requires custom container build (no public Dynamo+TokenSpeed image yet) and raw Deployments/Services instead of DynamoGraphDeployment (operator backend support pending). |
| DeepSeek-V4-Flash | vLLM | Aggregated | 4x B200 | ✅ | Text only — MoE model (284B / 13B active), DP=4 + EP, FP8 KV cache, reasoning + tool calling. Requires custom container build. |
| DeepSeek-V4-Flash | vLLM | Aggregated | 4x GB200 | ✅ | Text only — MoE model (284B / 13B active), TP=4 + EP, deep_gemm_mega_moe, FP8 KV cache, reasoning + tool calling (single NVL4 tray). Requires custom container build. |
| DeepSeek-V4-Flash | SGLang | Aggregated | 4x B200 | ✅ | Text only — MoE model (284B / 13B active), TP=4, MXFP4 MoE via FlashInfer, EAGLE MTP (3 steps / 4 draft tokens), reasoning + tool calling. Prebuilt image available; optional custom container build. |
| DeepSeek-V4-Pro | vLLM | Aggregated | 8x B200 | ✅ | Text only — MoE model (1.6T / 49B active, 1M context), TP=8 + EP, FP4+FP8 mixed checkpoint, FP8 KV cache, CSA+HCA attention, tool calling. Thinking modes unstable on Day-0 — run with thinking: false. Requires custom container build. |
| DeepSeek-V4-Pro | vLLM | Aggregated | 8x GB200 (2 NVL4 trays) | ✅ | Text only — same model as B200 agg; TP=8 + EP cross-node via NVLink72 (MNNVL) + ComputeDomain. Requires custom container build. |
| DeepSeek-V4-Pro | vLLM | Disaggregated | 16x GB200 (4 NVL4 trays) | ✅ | Text only — DP=8 + EP per worker, 1P + 1D, NVLink72 (MNNVL) + ComputeDomain. Requires custom container build. |
| DeepSeek-V4-Pro | SGLang | Aggregated | 8x B200 | ✅ | Text only — MoE model (1.6T / 49B active, 1M context), TP=8, MXFP4 MoE via FlashInfer, EAGLE MTP (3 steps / 4 draft tokens), reasoning + tool calling. Prebuilt image available (shared with DeepSeek-V4-Flash). |
| Solar-Open2-250B-NVFP4 | vLLM | Aggregated | 4x B200 | ✅ | Text only — NVFP4 checkpoint, 2 replicas × TP2, reasoning + tool calling. Prebuilt image available. |
| Solar-Open2-250B-NVFP4 | vLLM | Disaggregated | 8x B200 | ✅ | Text only — NVFP4 checkpoint, 1× prefill TP4 + EP and 2× decode TP2, NIXL/RDMA KV transfer, reasoning + tool calling. Prebuilt image available. |
Each complete recipe follows this standard structure:
<model-name>/
├── README.md (optional) # Model-specific deployment notes
├── model-cache/
│ ├── model-cache.yaml # PersistentVolumeClaim for model storage
│ └── model-download.yaml # Job to download model from HuggingFace
└── <framework>/ # vllm, sglang, or trtllm
└── <deployment-mode>/ # agg, disagg, disagg-single-node, etc.
├── deploy.yaml # Complete DynamoGraphDeployment manifest
└── perf.yaml (optional) # AIPerf benchmark job
In addition, accuracy/ is a shared, model-agnostic accuracy
check (deliberately outside the per-model structure above): point it at any
deployed recipe to compare the served model's benchmark score against its
model card. See accuracy/README.md.
1. Dynamo Platform Installed
The recipes require the Dynamo Kubernetes Platform to be installed. Follow the installation guide:
- Kubernetes Deployment Guide - Quickstart (~10 minutes)
- Detailed Installation Guide - Advanced options
2. GPU Cluster Requirements
Ensure your cluster has:
- GPU nodes matching recipe requirements (see table above)
- GPU operator installed
- Appropriate GPU drivers and container runtime
3. HuggingFace Access
Configure authentication to download models:
export NAMESPACE=your-namespace
kubectl create namespace ${NAMESPACE}
# Create HuggingFace token secret
kubectl create secret generic hf-token-secret \
--from-literal=HF_TOKEN="your-token-here" \
-n ${NAMESPACE}4. Storage Configuration
Update the storageClassName in <model>/model-cache/model-cache.yaml to match your cluster:
# Find your storage class name
kubectl get storageclass
# Edit the model-cache.yaml file and update:
# spec:
# storageClassName: "your-actual-storage-class"Step 1: Download Model
cd recipes
# Update storageClassName in model-cache.yaml first!
kubectl apply -f <model>/model-cache/ -n ${NAMESPACE}
# Wait for download to complete (may take 10-60 minutes depending on model size)
kubectl wait --for=condition=Complete job/model-download -n ${NAMESPACE} --timeout=6000s
# Monitor progress
kubectl logs -f job/model-download -n ${NAMESPACE}Step 2: Deploy Service
Update the image in <model>/<framework>/<mode>/deploy.yaml.
kubectl apply -f <model>/<framework>/<mode>/deploy.yaml -n ${NAMESPACE}
# Check deployment status
kubectl get dynamographdeployment -n ${NAMESPACE}
# Check pod status
kubectl get pods -n ${NAMESPACE}
# Wait for pods to be ready
kubectl wait --for=condition=ready pod -l nvidia.com/dynamo-graph-deployment-name=<deployment-name> -n ${NAMESPACE} --timeout=600sStep 3: Test Deployment
# Port forward to access the service locally
kubectl port-forward svc/<deployment-name>-frontend 8000:8000 -n ${NAMESPACE}
# In another terminal, test the endpoint
curl http://localhost:8000/v1/models
# Send a test request
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<model-name>",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 50
}'Step 4: Run Benchmark (Optional)
# Only if perf.yaml exists in the recipe directory
kubectl apply -f <model>/<framework>/<mode>/perf.yaml -n ${NAMESPACE}
# Monitor benchmark progress
kubectl logs -f job/<benchmark-job-name> -n ${NAMESPACE}
# View results after completion
kubectl logs job/<benchmark-job-name> -n ${NAMESPACE} | tail -50export NAMESPACE=dynamo-demo
kubectl create namespace ${NAMESPACE}
# Create HF token secret
kubectl create secret generic hf-token-secret \
--from-literal=HF_TOKEN="your-token" \
-n ${NAMESPACE}
# Deploy
cd recipes
kubectl apply -f qwen3-32b-fp8/model-cache/ -n ${NAMESPACE}
kubectl wait --for=condition=Complete job/model-download -n ${NAMESPACE} --timeout=6000s
kubectl apply -f qwen3-32b-fp8/trtllm/agg/deploy.yaml -n ${NAMESPACE}
# Test
kubectl port-forward svc/qwen3-32b-fp8-agg-frontend 8000:8000 -n ${NAMESPACE}For Qwen3-0.6B with vLLM (Aggregated), an example of integration with the Inference Gateway is provided.
First, deploy the Dynamo Graph per instructions above, substituting the Qwen3-0.6B recipe.
Then follow Deploy Inference Gateway Section 2 to install GAIE.
Update the containers.epp.image in the deployment file, i.e. qwen3-0.6b/vllm/agg/gaie/deploy.yaml. It should match the release tag and be in the format nvcr.io/nvidia/ai-dynamo/frontend:<version> e.g. nvcr.io/nvidia/ai-dynamo/frontend:0.9.0
The recipe assumes you are using Kubernetes discovery backend and sets the DYN_DISCOVERY_BACKEND env variable in the epp deployment. If you want to use etcd enable the lines below and remove the DYN_DISCOVERY_BACKEND env var.
- name: ETCD_ENDPOINTS
value: "dynamo-platform-etcd.$(PLATFORM_NAMESPACE):2379" # update dynamo-platform to appropriate namespaceexport DEPLOY_PATH=qwen3-0.6b/vllm/agg/
# DEPLOY_PATH=<model>/<framework>/<mode>/
kubectl apply -R -f "$DEPLOY_PATH/gaie" -n "$NAMESPACE"See deepseek-r1/trtllm/disagg/wide_ep/gb200/deploy.yaml for the complete multi-node WideEP configuration.
Each deploy.yaml contains:
- ConfigMap: Engine-specific configuration (embedded in the manifest)
- DynamoGraphDeployment: Kubernetes resource definitions
- Resource limits: GPU count, memory, CPU requests/limits
- Image references: Container images with version tags
Model Configuration:
# In deploy.yaml under worker args:
args:
- python3 -m dynamo.vllm --model <your-model-path> --served-model-name <name>GPU Resources:
resources:
limits:
gpu: "4" # Adjust based on your requirements
requests:
gpu: "4"Scaling:
services:
decode:
replicas: 2 # Scale to multiple workersRouter Mode:
# In Frontend args:
args:
- python3 -m dynamo.frontend --router-mode kv --http-port 8000
# Options: round-robin, kv (KV-aware routing)Container Images:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:x.y.z
# Update version tag as neededPods stuck in Pending:
- Check GPU availability:
kubectl describe node <node-name> - Verify storage class exists:
kubectl get storageclass - Check resource requests vs. available resources
Model download fails:
- Verify HuggingFace token is correct
- Check network connectivity from cluster
- Review job logs:
kubectl logs job/model-download -n ${NAMESPACE}
Workers fail to start:
- Check GPU compatibility (driver version, CUDA version)
- Verify image pull secrets if using private registries
- Review pod logs:
kubectl logs <pod-name> -n ${NAMESPACE}
For more troubleshooting:
- Kubernetes Deployment Guide - Platform installation and concepts
- API Reference - DynamoGraphDeployment CRD specification
- vLLM Backend Guide - vLLM-specific features
- SGLang Backend Guide - SGLang-specific features
- TensorRT-LLM Backend Guide - TensorRT-LLM features
- Observability - Monitoring and logging
- Benchmarking Guide - Performance testing
We welcome contributions of new recipes! See CONTRIBUTING.md for:
- Recipe submission guidelines
- Required components checklist
- Testing and validation requirements
- Documentation standards
A production-ready recipe must include:
- ✅ Complete
deploy.yamlwith DynamoGraphDeployment - ✅ Model cache PVC and download job
- ✅ Benchmark recipe (
perf.yaml) for performance testing - ✅ Verification on target hardware
- ✅ Documentation of GPU requirements