| title | Google Kubernetes Engine (GKE) |
|---|
gcloud container clusters create ${CLUSTER_NAME}
--project=${PROJECT_ID}
--location=${ZONE}
--subnetwork=default
--disk-size=${DISK_SIZE}
--machine-type=${CLUSTER_MACHINE_TYPE}
--num-nodes=${CPU_NODE}
#### Create GPU pool
```bash
gcloud container node-pools create gpu-pool \
--accelerator type=${GPU_TYPE},count=${GPU_COUNT},gpu-driver-version=latest \
--project=${PROJECT_ID} \
--location=${ZONE} \
--cluster=${CLUSTER_NAME} \
--machine-type=${NODE_POOL_MACHINE_TYPE} \
--disk-size=${DISK_SIZE} \
--num-nodes=${GPU_NODE} \
--enable-autoscaling \
--min-nodes=1 \
--max-nodes=3
git clone https://github.com/ai-dynamo/dynamo.git
cd dynamo
# Checkout the release branch matching your Dynamo platform version
git checkout release/1.5.0export HF_TOKEN=<HF_TOKEN>
kubectl create secret generic hf-token-secret
--from-literal=HF_TOKEN=${HF_TOKEN}
-n ${NAMESPACE}
</Step>
<Step title="Install Dynamo Kubernetes Platform" id="install-dynamo-kubernetes-platform">
[See installation steps](../../install-dynamo.md#overview)
After installation, verify the installation:
**Expected output**
```bash
kubectl get pods
NAME READY STATUS RESTARTS AGE
dynamo-platform-dynamo-operator-controller-manager-69b9794fpgv9 2/2 Running 0 4m27s
dynamo-platform-etcd-0 1/1 Running 0 4m27s
dynamo-platform-nats-0 2/2 Running 0 4m27s
In the deployment yaml file, some adjustments have to/ could be made:
- (Required) Add args to change
LD_LIBRARY_PATHandPATHof decoder container, to enable GKE find the correct GPU driver - Change VLLM image to the desired one on NGC
- Add namespace to metadata
- Adjust GPU/CPU request and limits
- Change model to deploy
More configurations please refer to https://github.com/ai-dynamo/dynamo/tree/main/examples/deployments/GKE/vllm
Please note that LD_LIBRARY_PATH needs to be set properly in GKE as per Run GPUs in GKE
The following snippet needs to be present in the args field of the deployment yaml file:
export LD_LIBRARY_PATH=/usr/local/nvidia/lib64:$LD_LIBRARY_PATH
export PATH=$PATH:/usr/local/nvidia/bin:/usr/local/nvidia/lib64
/sbin/ldconfigFor example, refer to the release/1.5.0 disaggregated manifest:
metadata:
name: vllm-disagg
spec:
components:
- name: decode
podTemplate:
spec:
containers:
- args:
- |
export LD_LIBRARY_PATH=/usr/local/nvidia/lib64:$LD_LIBRARY_PATH
export PATH=$PATH:/usr/local/nvidia/bin:/usr/local/nvidia/lib64
/sbin/ldconfig
python3 -m dynamo.vllm --model Qwen/Qwen3-0.6B --disaggregation-mode decode
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0
name: main
resources:
limits:
nvidia.com/gpu: "1"kubectl apply -f disagg.yaml -n ${NAMESPACE}
**Expected output after successful deployment**
```bash
kubectl get pods
NAME READY STATUS RESTARTS AGE
dynamo-platform-dynamo-operator-controller-manager-c665684ssqkx 2/2 Running 0 65m
dynamo-platform-etcd-0 1/1 Running 0 65m
dynamo-platform-nats-0 2/2 Running 0 65m
vllm-disagg-frontend-5954ddc4dd-4w2cb 1/1 Running 0 11m
vllm-disagg-decode-77844cfcff-ddn4v 1/1 Running 0 11m
vllm-disagg-prefill-55d5b74b4f-zrskh 1/1 Running 0 11m
export FRONTEND_POD=$(kubectl get pods -n ${NAMESPACE} | grep "${DEPLOYMENT_NAME}-frontend" | sort -k1 | tail -n1 | awk '{print $1}')
kubectl port-forward deployment/vllm-disagg-frontend 8000:8000 -n ${NAMESPACE}
curl localhost:8000/v1/chat/completions
-H "Content-Type: application/json"
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [
{
"role": "user",
"content": "In the heart of Eldoria, an ancient land of boundless magic and mysterious creatures, lies the long-forgotten city of Aeloria. Once a beacon of knowledge and power, Aeloria was buried beneath the shifting sands of time, lost to the world for centuries. You are an intrepid explorer, known for your unparalleled curiosity and courage, who has stumbled upon an ancient map hinting at ests that Aeloria holds a secret so profound that it has the potential to reshape the very fabric of reality. Your journey will take you through treacherous deserts, enchanted forests, and across perilous mountain ranges. Your Task: Character Background: Develop a detailed background for your character. Describe their motivations for seeking out Aeloria, their skills and weaknesses, and any personal connections to the ancient city or its legends. Are they driven by a quest for knowledge, a search for lost familt clue is hidden."
}
],
"stream":false,
"max_tokens": 30
}'
### Response
```json
{"id":"chatcmpl-bd0670d9-0342-4eea-97c1-99b69f1f931f","choices":[{"index":0,"message":{"content":"Okay, here's a detailed character background for your intrepid explorer, tailored to fit the premise of Aeloria, with a focus on a","refusal":null,"tool_calls":null,"role":"assistant","function_call":null,"audio":null},"finish_reason":"stop","logprobs":null}],"created":1756336263,"model":"Qwen/Qwen3-0.6B","service_tier":null,"system_fingerprint":null,"object":"chat.completion","usage":{"prompt_tokens":190,"completion_tokens":29,"total_tokens":219,"prompt_tokens_details":null,"completion_tokens_details":null}}