Full-parameter GRPO of Inkling-Small (Thinking Machines Lab; 276B-parameter MoE with about 12B active, 42 layers) with Miles. Every rollout is one terminal task: the Harbor agent server starts the task's sandbox (Daytona, Modal, GKE or Docker), runs the agent against the policy and verifies the result; the verifier's score is the reward. Two agents are supported, one script each:
| script | agent | how the model's output is read |
|---|---|---|
run_harbor_camel.sh |
Harbor's CAMEL agent, Inkling's native tool calls | the Miles session server parses Inkling's completions; SGLang's parsers are off |
run_harbor_terminus2.sh |
Harbor's Terminus-2 agent, JSON commands in plain text, interleaved thinking | SGLang's Inkling tool and reasoning parsers split the output; needs a patched SGLang (Step 1) |
train.py |
the Miles launcher both scripts call (python train.py train --help) |
|
prepare_model.py |
downloads the model and converts it for Megatron (Step 3) |
Everything else (parallelism, serving, GRPO settings, batch shape) is the same for both.
Inkling-Small is released under the Apache-2.0 license together with the Thinking Machines Model Acceptable Use Policy; your use of the model and of anything you train from it must follow both.
- GPUs: 8 nodes x 8 H200 (141 GB). Training and rollout share every GPU. The actor runs
tensor parallel 8 with sequence parallel, expert parallel 8 and one pipeline stage per node
(TP8/PP8/EP8, no context or data parallelism); each node also serves one TP8 SGLang engine
(8 engines). 6 or 7 nodes also work (
NUM_NODES, 6 or 7 pipeline stages) with less memory headroom. - Host memory: the training weights and the Adam state are offloaded to CPU memory while SGLang serves, so plan for several hundred GB of free RAM per node.
- Disk, on storage every node can read at the same path: the Hugging Face checkpoint and its
Megatron conversion (about 550 GB each: 276B parameters in BF16), plus about 550 GB per saved
checkpoint under
RUNS_ROOT. - Head node: also runs the Harbor agent server and the Miles session servers (CPU).
- Accounts: Hugging Face (the model is not gated), one sandbox provider (Daytona, Modal or your own GKE cluster; or a local Docker daemon), optionally Weights & Biases.
| component | CAMEL (run_harbor_camel.sh) |
Terminus-2 (run_harbor_terminus2.sh) |
|---|---|---|
| Docker image | radixark/miles@sha256:946f29396ac313d03b50ac0c0c8930b7eaaa403ae9e7d6115595d57e59cdad12 |
same |
| Miles | Michaelsqj/miles@38bef2a605191c137323976c211ebdac467b3295: the session server parses raw Inkling completions (tool calls, reasoning), so SGLang's parsers can stay off (later merged upstream as radixark/miles#2302); session servers get 180 s to start |
Michaelsqj/miles@80a25cb568982b8e445498ab7d3a0fc3d5d3670e: keeps Inkling's ordered thinking/text blocks when the agent replays its history, so the replayed turns re-tokenize to the generated tokens; session-server startup time scales with the pool size |
| SGLang | sgl-project/sglang@cb05a44f35a7c9e27e46d74112cc841ca674ef43 (sglang-miles branch) |
michaelsqj/sglang@baccf651fe984a825c35b2554db3572b31e9e25a: cb05a44f plus an ordered content_blocks field in chat responses for Inkling's thinking/text blocks (not merged upstream) |
| Harbor | Michaelsqj/harbor@acac1c20e0350f70c60fc6a0755a99d9302b90dd: CAMEL's response_feedback switches reach the agent, CAMEL library logging off; pins CAMEL to Michaelsqj/camel@1c42729b29b6f9f216bab2d08a5e7a52bbb6cf8e in its uv.lock |
Michaelsqj/harbor@1e02f96f00a18b5e9158a4baffbaa2bb5a0b2146: agent server with one subprocess worker per trial |
| Model | thinkingmachines/Inkling-Small@8cc5877b44d343f88b92086aa1fb72897950f06a |
same |
HARBOR_COMMIT is installed automatically (Step 4). The scripts warn when Miles or SGLang (head
node) is not at the commit above; STRICT_PINS=1 turns that into an error. They refuse to start
when a feature they depend on is missing from the Miles, SGLang or Harbor checkout (see Notes).
docker pull radixark/miles@sha256:946f29396ac313d03b50ac0c0c8930b7eaaa403ae9e7d6115595d57e59cdad12
docker run -d --name miles --gpus all --network host --ipc host --shm-size 64g \
--ulimit memlock=-1 --ulimit stack=67108864 \
-v /path/to/models:/root/models \
-v /path/to/data:/data \
-v /path/to/seta:/workspace/seta \
radixark/miles@sha256:946f29396ac313d03b50ac0c0c8930b7eaaa403ae9e7d6115595d57e59cdad12 sleep infinity
docker exec -it miles bashAdd your cluster's InfiniBand/RDMA device flags if it needs them, and
-v /var/run/docker.sock:/var/run/docker.sock on the head node for SANDBOX_BACKEND=docker.
Inside the container on every node, check out the Miles and SGLang commits of the script you
will run. Miles runs from /root/miles (the scripts put it on PYTHONPATH) and SGLang is
imported from /sgl-workspace/sglang, so a checkout is all it takes:
# run_harbor_camel.sh
git -C /root/miles fetch https://github.com/Michaelsqj/miles.git 38bef2a605191c137323976c211ebdac467b3295
git -C /root/miles checkout --detach 38bef2a605191c137323976c211ebdac467b3295
git -C /sgl-workspace/sglang fetch https://github.com/sgl-project/sglang.git cb05a44f35a7c9e27e46d74112cc841ca674ef43
git -C /sgl-workspace/sglang checkout --detach cb05a44f35a7c9e27e46d74112cc841ca674ef43
# run_harbor_terminus2.sh
git -C /root/miles fetch https://github.com/Michaelsqj/miles.git 80a25cb568982b8e445498ab7d3a0fc3d5d3670e
git -C /root/miles checkout --detach 80a25cb568982b8e445498ab7d3a0fc3d5d3670e
git -C /sgl-workspace/sglang fetch https://github.com/michaelsqj/sglang.git baccf651fe984a825c35b2554db3572b31e9e25a
git -C /sgl-workspace/sglang checkout --detach baccf651fe984a825c35b2554db3572b31e9e25a
# check: SGLang is imported from the checkout
python -c "import os, sglang; print(os.path.dirname(sglang.__file__))" # /sgl-workspace/sglang/python/sglangIf the check prints another path, install the checkout with
pip install --no-deps -e /sgl-workspace/sglang/python. The Terminus-2 SGLang commit changes
only Python files on top of the CAMEL one, so no kernels need rebuilding.
To switch between the two scripts later, stop the running job and repeat this on every node.
# head node
ray start --head --node-ip-address "$HEAD_IP" --port 6379 \
--dashboard-host 0.0.0.0 --dashboard-port 8265 --num-gpus 8 --disable-usage-stats
# every other node
ray start --address "$HEAD_IP:6379" --node-ip-address "$NODE_IP" --num-gpus 8 --disable-usage-stats
# head node: expect 64 GPUs
ray statusThe run scripts connect to this cluster (they never start or stop Ray) and must run on the head
node, inside the container. Before submitting, Miles kills leftover sglang and miles
processes on the head node.
On one node, inside the container, with 8 free GPUs:
cd /workspace/seta
PYTHONPATH=/root/miles python scripts/miles/examples/inkling_grpo/prepare_model.py --model-root /root/modelsThis downloads thinkingmachines/Inkling-Small at the tested revision with the hf CLI
(pip install -U huggingface_hub if it is missing) and converts it to Megatron's torch_dist
format (TP8/EP8; the actor re-shards it on load). Both steps can be re-run and resume. Output:
/root/models/Inkling-Small/ Hugging Face checkpoint (SGLang serves it; Miles reads its config)
/root/models/Inkling-Small_torch_dist/ Megatron checkpoint (actor start weights and KL reference)
Every node must see both directories at the same path (shared storage, or copy them).
Pick one with SANDBOX_BACKEND (in .env, Step 6). Trials run in sandboxes of that backend;
the agent itself runs in the Harbor agent server on the head node.
The default. Set DAYTONA_API_KEY (https://app.daytona.io -> API keys), and DAYTONA_API_URL
for a non-default endpoint.
Set MODAL_TOKEN_ID / MODAL_TOKEN_SECRET, or run modal token new in the container on the head
node (writes ~/.modal.toml).
Your own Google Cloud project. Only the sandboxes run in GKE; the agent server stays on the head node and Cloud Build builds each task's image into Artifact Registry the first time the task runs.
Prerequisites. A project with billing, and the gcloud CLI
with gcloud components install gke-gcloud-auth-plugin kubectl, logged in (gcloud auth login, or
gcloud auth activate-service-account --key-file=...) where you run the commands below and inside
the container on the head node, where the agent server runs. create enables the Kubernetes Engine,
Compute Engine, Artifact Registry, Cloud Build and IAM APIs. IAM: Owner, or the roles listed at the top
of ../../common/gke_cluster.sh (to set up: Kubernetes Engine Admin, Artifact Registry Admin, Service
Account Admin and User, Project IAM Admin, Service Usage Admin; to train: Kubernetes Engine Developer,
Artifact Registry Reader, Cloud Build Editor, Storage Admin).
Put SANDBOX_BACKEND=gke, GKE_PROJECT_ID, GKE_REGION and GKE_CLUSTER_NAME in .env (Step 6); the
GKE scripts read them from ./.env too. Then, from this folder:
bash ../../common/gke_cluster.sh create # registry, node service account, cluster, sandbox pool at 0 nodes
HARBOR_COMMIT=<HARBOR_COMMIT of the run script> bash ../../common/gke_tools_image.sh # prints GKE_TOOLS_IMAGE=...: add it to .env
bash ../../common/gke_cluster.sh apply-network-policies # optional: sandboxes get DNS and web only (--no-web-egress: no internet)
bash ../../common/gke_cluster.sh scale 8 # before a run: 8 n2-standard-16 sandbox nodes
bash ../../common/gke_cluster.sh down --yes # after the run: sandbox nodes to zeroThe launcher writes the cluster's kubeconfig on first use (gke_cluster.sh credentials, to
~/.kube/gke_<project>_<region>_<cluster>). An n2-standard-16 node holds about 15 one-CPU sandboxes;
scale for AGENT_MAX_CONCURRENT. status shows nodes and sandbox pods; clear-sandboxes --yes deletes
pods left by a crashed run. DRY_RUN=1 prints every command instead of running it. For an existing
cluster skip create (network policies need Dataplane V2) and set GKE_REGISTRY_NAME to a Docker
repository in GKE_REGION. Cost: sandbox nodes bill while they are up, so run down --yes after
every run; the control plane and one small system node per zone bill until gke_cluster.sh delete --yes.
The first run on a new task set also pays for Cloud Build time and registry storage.
Uses the Docker daemon of the head node: mount its socket into the container (Step 1). Only
practical with a small AGENT_MAX_CONCURRENT.
On GKE these scripts size every sandbox at 1 CPU, 2048 MiB memory and 6144 MiB storage
(SANDBOX_CPUS, SANDBOX_MEMORY_MB, SANDBOX_STORAGE_MB), the size they were run with; size
the cluster for AGENT_MAX_CONCURRENT (160) sandboxes. The other backends use each task's own
resources.
Harbor is installed by the run script on first use, at HARBOR_COMMIT, into its own virtualenv
under ~/.cache/seta/harbor/<commit> (HARBOR_HOME) on the head node. It needs Python 3.12 (uv
fetches it when installed: pip install uv). To install ahead of time:
HARBOR_COMMIT=acac1c20e0350f70c60fc6a0755a99d9302b90dd bash scripts/miles/common/harbor_install.sh # CAMEL
HARBOR_COMMIT=1e02f96f00a18b5e9158a4baffbaa2bb5a0b2146 bash scripts/miles/common/harbor_install.sh # Terminus-2No data ships with this recipe. It needs two inputs, described in
../../data/README.md:
PROMPT_DATA: a JSONL with one row per task (metadata.instance_idnames the task).TASKS_DIR: the Harbor task directories those rows name (<task>/task.toml,instruction.md,environment/,tests/).
The example files in scripts/miles/data are Terminal-Bench examples for smoke tests only; do not
train on them. Train on your own task set, for example the
SETA-Env environments:
# download the tasks as described in ../../data/README.md, e.g. into dataset/seta-env-final/<task>/
python scripts/miles/common/build_prompt_dataset.py build \
--tasks-dir dataset/seta-env-final --agent-name camel --output train.jsonlmetadata.agent_name is informational; the script decides the agent. The scripts check that
every row has a task directory before they start. GRPO learns only from prompt groups whose
samples get different rewards, so tasks the model sometimes solves are the useful ones.
cd scripts/miles/examples/inkling_grpo
cp env.example .env # then edit .envBoth scripts read .env from this folder (or the file named by ENV_FILE); values there and in
the environment override the script defaults.
| variable | required | meaning |
|---|---|---|
HEAD_IP |
yes (or CLUSTER_CONFIG) |
Ray head node IP; the scripts run on this node |
CLUSTER_CONFIG |
a copy of ../../cluster.example.yaml with your IPs, instead of HEAD_IP |
|
NCCL_SOCKET_IFNAME, GLOO_SOCKET_IFNAME |
interconnect interface, when it is not the default route | |
MODEL_ROOT |
yes | where prepare_model.py wrote the model (default /root/models) |
MODEL_DIR, TORCH_DIST_DIR |
the two checkpoints, if not under MODEL_ROOT |
|
LOAD_DIR |
start from an earlier run's checkpoints/ (see Outputs & resume) |
|
PROMPT_DATA, TASKS_DIR |
yes | Step 5 |
SANDBOX_BACKEND |
yes | daytona (default), modal, gke or docker |
DAYTONA_API_KEY, DAYTONA_API_URL |
daytona | Step 4 |
MODAL_TOKEN_ID, MODAL_TOKEN_SECRET |
modal | Step 4 (or ~/.modal.toml) |
GKE_* |
gke | Step 4 |
WANDB_API_KEY |
turns W&B logging on; passed to Miles as --wandb-key |
|
WANDB_PROJECT, WANDB_ENTITY |
W&B project (default seta) and team (default: your user) |
|
RUNS_ROOT, RUN_NAME |
outputs go to $RUNS_ROOT/$RUN_NAME (default ./runs/inkling-grpo-<agent>-<UTC time>); RUNS_ROOT must be shared storage, every node writes checkpoint shards there |
|
TORCHINDUCTOR_CACHE_DIR |
compile cache, default $RUNS_ROOT/torchinductor_cache; keep it across runs |
|
MILES_DIR, SGLANG_DIR, MEGATRON_PATH |
checkouts in the container (defaults /root/miles, /sgl-workspace/sglang, /root/Megatron-LM) |
|
STRICT_PINS |
1: fail instead of warn when Miles/SGLang are not at the tested commits |
|
HARBOR_COMMIT, HARBOR_HOME |
Harbor commit and install location (defaults: the script's pin, ~/.cache/seta/harbor) |
The training settings are listed in the Configuration reference.
Run on the head node inside the container, in tmux or screen (a run takes many hours).
DRY_RUN=1 prints the resolved train.py command and exits before anything is installed or
started.
cd /workspace/seta
DRY_RUN=1 bash scripts/miles/examples/inkling_grpo/run_harbor_camel.sh
bash scripts/miles/examples/inkling_grpo/run_harbor_camel.shNeeds the Miles and SGLang commits of this script on every node (Step 1).
cd /workspace/seta
DRY_RUN=1 bash scripts/miles/examples/inkling_grpo/run_harbor_terminus2.sh
bash scripts/miles/examples/inkling_grpo/run_harbor_terminus2.shEach script checks the model, the checkouts and the prompt data, installs Harbor, starts the Harbor agent server on the head node (stopped again when the script exits) and submits the Miles job to Ray. A quick end-to-end check with a small batch:
NUM_ROLLOUT=3 ROLLOUT_BATCH_SIZE=4 N_SAMPLES_PER_PROMPT=2 OVER_SAMPLING_BATCH_SIZE=6 \
bash scripts/miles/examples/inkling_grpo/run_harbor_camel.sh(GLOBAL_BATCH_SIZE and AGENT_MAX_CONCURRENT follow the batch shape.)
$RUNS_ROOT/$RUN_NAME/logs/train.log: the Miles job (rollout progress, rewards, losses).$RUNS_ROOT/$RUN_NAME/logs/agent_server.log: the Harbor agent server.$RUNS_ROOT/$RUN_NAME/trials/: one folder per trial (agent log, trajectory, verifier output).- Ray dashboard:
http://<HEAD_IP>:8265. - Harbor dashboard:
http://<HEAD_IP>:11000/dashboard/(AGENT_SERVER_PORT). - W&B, when
WANDB_API_KEYis set: KL to the reference model and the entropy are logged every step (the KL coefficient is 0, so they are only observed).
$RUNS_ROOT/$RUN_NAME/
logs/train.log, logs/agent_server.log
trials/ Harbor trial artifacts
checkpoints/ Megatron torch_dist checkpoints, weights only
dump_details/ per-rollout samples and token ids (DUMP_DETAILS=1)
wandb/
run_manifest.md commits, backend, prompt data (row count, sha256)
launcher/ copies of the run script and train.py
$RUNS_ROOT/torchinductor_cache/ compile cache shared by runs
A checkpoint is written every SAVE_INTERVAL rollouts and after the last one. Checkpoints hold
the model weights only (with the Adam state a checkpoint of this model is several TB), so a run
cannot be resumed exactly. To continue from saved weights, start a new run with
LOAD_DIR=$RUNS_ROOT/<old run>/checkpoints: the actor loads the latest checkpoint there, the
optimizer state and the rollout count start fresh, and the KL reference stays the base model.
Miles' tools/convert_torch_dist_to_hf.py converts a checkpoint to Hugging Face format; these
scripts do not wrap it.
Environment variables (or .env) read by both scripts; defaults are the configuration these
recipes were run with.
| variable | default | meaning |
|---|---|---|
NUM_NODES |
8 |
nodes of 8 GPUs; 6, 7 or 8 (one pipeline stage and one SGLang engine per node) |
GPUS_PER_NODE |
8 |
only 8 is supported |
LR |
1e-5 |
constant learning rate (Adam, weight decay 0.1) |
NUM_ROLLOUT |
100 |
rollouts, one optimizer step each |
OVER_SAMPLING_BATCH_SIZE |
20 |
prompt groups started per rollout |
ROLLOUT_BATCH_SIZE |
16 |
complete groups kept per rollout (groups with an aborted sample are dropped) |
N_SAMPLES_PER_PROMPT |
8 |
trials per prompt (the GRPO group) |
GLOBAL_BATCH_SIZE |
128 |
ROLLOUT_BATCH_SIZE x N_SAMPLES_PER_PROMPT: one step per rollout |
SAVE_INTERVAL |
99 (CAMEL), 50 (Terminus-2) |
checkpoint every N rollouts (plus the last) |
MAX_SEQ_LEN |
32768 |
tokens of one training sample (prompt plus all turns); longer samples are truncated |
MAX_RESPONSE_LEN |
16384 |
tokens of one model turn |
SESSION_SERVERS |
64 |
Miles session-server processes on the head node |
SESSION_SERVER_PORT |
30000 |
first port; they use [port, port + SESSION_SERVERS) |
DUMP_DETAILS |
1 |
write dump_details/ |
EXTRA_ARGS |
extra Miles/Megatron/SGLang flags, appended last (the last occurrence of a flag wins) | |
AGENT_MAX_CONCURRENT |
160 |
trials Harbor runs at once; default OVER_SAMPLING_BATCH_SIZE x N_SAMPLES_PER_PROMPT |
AGENT_SERVER_PORT |
11000 |
Harbor agent server port on the head node |
HARBOR_AGENT_MAX_ITERATIONS |
50 |
agent turns per trial |
HARBOR_MAX_SEQ_LEN |
1048576 |
the agent's context bound (separate from MAX_SEQ_LEN) |
HARBOR_AGENT_CALL_TIMEOUT_SEC |
10800 |
client deadline for one whole trial (queue, sandbox, agent, verification) |
HARBOR_AGENT_TIMEOUT_MULTIPLIER |
12 |
scales each task's own agent timeout |
AGENT_MODEL_NAME |
inkling-small |
model name the agent sends |
SANDBOX_CPUS |
1 |
GKE only: CPUs per sandbox (other backends use the task's value) |
SANDBOX_MEMORY_MB |
2048 |
GKE only: memory per sandbox |
SANDBOX_STORAGE_MB |
6144 |
GKE only: ephemeral storage per sandbox |
HARBOR_INTERLEAVED_THINKING |
true |
Terminus-2 only: earlier turns keep their reasoning |
Fixed in train.py (edit it, or override with EXTRA_ARGS): GRPO with clip 0.2/0.28 and
truncated importance sampling, rollout routing replay for the MoE layers, rollout temperature 1,
full activation recomputation, micro-batch 1, and SGLang at memory fraction 0.75, 256 running
requests and 1,048,576 KV tokens per engine.
Not available at these Harbor commits: HARBOR_TERMINUS_PARSER,
HARBOR_TERMINUS_ENABLE_SUMMARIZE, HARBOR_CAMEL_MAX_COMPACTIONS (the scripts unset them),
HARBOR_AGENT_TIMEOUT_SEC, HARBOR_EXTRA_ARTIFACTS and HARBOR_EXTRA_COLLECT (leave unset).
- Learning rate. The default is
1e-5. For full fine-tuning of this model3e-5was the highest learning rate that trained stably;5e-5and1e-4collapsed. Watch the logged KL to the reference and the entropy even though neither is in the loss, and watch for the policy learning to end hard tasks early. - Serving on the training GPUs. One TP8/EP1 engine per node at memory fraction 0.75. TP4 engines leave no room for the weight update after each step, and SGLang expert parallelism does not shrink the per-GPU footprint. The large KV pool (1M tokens per engine) avoids request retractions with 160 concurrent multi-turn trials. The router keeps each session on one engine (prefix cache) and places new sessions on the least-loaded engine.
- Training mesh and memory. TP8/PP8/EP8 without context parallelism, with
MAX_SEQ_LENcapping each sample: without the cap, a single long trajectory can run the MoE layers out of memory. Context parallel 2 with all-gather (--context-parallel-size 2 --allgather-cp) worked in offline replays but hit NCCL timeouts in live runs on H200; avoid it. The offload flags intrain.py(--offload-train-target cpu --optimizer-cpu-offload --overlap-cpu-optimizer-d2h-h2d --use-precision-aware-optimizer) are required: without them creating the Adam state runs out of GPU memory. KeepTORCHINDUCTOR_CACHE_DIRacross runs: with a cold cache the first actor step spends a long time compiling. - Storage. Checkpoints are weights only (
--no-save-optim); keepSAVE_INTERVALsparse andRUNS_ROOTon a large volume.DUMP_DETAILS=1and the trial folders also grow with every rollout. - Session servers. 64 session-server processes start at once and each imports the tokenizer stack; Miles' old 30 s readiness deadline rejects the pool before the first rollout. Both Miles commits give them longer, and the scripts refuse a Miles with the 30 s deadline.
- Token-in/token-out. Training uses the exact tokens the engine produced, so the agent's
replayed history must re-tokenize to them. CAMEL: SGLang's Inkling parsers must stay off
(
train.pyrefuses them) because the Miles session server parses Inkling's completions. Terminus-2: the SGLang and Miles commits above keep Inkling's interleaved thinking/text blocks in order; the scripts refuse checkouts without them. - Terminus-2 protocol. Inkling sometimes emits its native tool-call blocks inside Terminus-2's
JSON protocol. When such trajectories still succeed they get a positive advantage and the
habit is reinforced; at
5e-5this degraded the Terminus-2 run quickly. Check the trials for native tool blocks, and prefer the CAMEL script, which uses the native format. - CAMEL behavior. A malformed or length-cut turn ends the trajectory as the policy wrote it: no corrective feedback, no retry. The Harbor commit routes these switches to the agent (with older commits they are silently ignored), and it turns off CAMEL's library logging, which otherwise logs every request payload and can fill the run disk. The script checks both.
- Head-node processes. Miles kills leftover
sglangandmilesprocesses on the head node before it submits a job; do not share the head node with another Miles job.