Skip to content

Latest commit

Β 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

SETA: Scaling Environments for Terminal Agents

SETA

Designing resilient toolkits and scalable RL environments for CAMEL terminal agents

πŸ€— Dataset πŸ€— Model arXiv Paper SETA Blog


Installation

git clone --recurse-submodules https://github.com/camel-ai/seta.git
cd seta
bash setup.sh

Quick start

Three runtime options β€” choose one:

Option A: Local Docker (single machine, no extra setup)

# uses eval_default.yaml (env_type: docker)
--config scripts/evaluation/configs/eval_default.yaml

Option B: Remote Docker (multiple nodes via slot pool service)

# 1. start slot pool service first
bash seta_env/runtimes/slot_pool_service/start.sh --dataset seta-env-v2
# 2. uses eval_remote.yaml (env_type: remote_docker)
--config scripts/evaluation/configs/eval_remote.yaml

Option C: Env Service (remote CPU servers for agent execution, see env_service)

# 1. deploy env_service to CPU servers + start scheduler
GH_TOKEN=ghp_xxx HF_TOKEN=hf_xxx bash seta_env/services/start.sh --dataset seta-env-v2
# 2. run eval via AReaL launcher
python -m areal.launcher.local scripts/areal/eval_env_service.py \
    --config scripts/areal/configs/config_eval_env_service_seta_v2.yaml

Evaluation

# start model server
python -m sglang.launch_server --model Qwen/Qwen3-8B --port 30000

# run eval (dataset auto-downloads on first use)
python scripts/evaluation/eval.py --config scripts/evaluation/configs/eval_default.yaml

# sweep across models and datasets
python scripts/evaluation/sweep_eval.py scripts/evaluation/configs/sweep.yaml

# results β†’ outputs/eval/<experiment>/<trial>/summary.json, results.csv

Training (AReaL)

# RL training
python -m areal.launcher.local \
    scripts/areal/rl_train.py \
    --config scripts/areal/configs/config_eval.yaml

# eval only (no gradient updates, single GPU)
python -m areal.launcher.local \
    scripts/areal/eval.py \
    --config scripts/areal/configs/config_eval.yaml \
    allocation_mode=sglang:d1p1t1+eval

# results β†’ outputs/areal/experiments/<experiment>/<trial>/

Training (Miles)

RL training of terminal agents with the Miles framework. Rollouts run through the Harbor agent server (Harbor's Terminus-2 or CAMEL agent in Daytona, Modal, GKE or Docker sandboxes) or through the seta env_service. Each example folder has a step-by-step README covering the container, Ray cluster, model preparation, sandboxes, dataset and launch:

Example Model Algorithm
deepseek_v4_grpo DeepSeek-V4-Flash GRPO
glm47_flash_grpo GLM-4.7-Flash GRPO
glm5_2_lora_grpo GLM-5.2 LoRA GRPO
glm5_2_lora_ppo GLM-5.2 LoRA PPO
inkling_grpo Inkling-Small GRPO
qwen3_8_27b_grpo Qwen3.8-27B GRPO
qwen3_8_27b_ppo Qwen3.8-27B PPO
cd scripts/miles/examples/qwen3_8_27b_grpo
cp env.example .env            # fill in cluster, model, data, sandbox and W&B settings
DRY_RUN=1 bash run_harbor_terminus2.sh   # print the resolved command
bash run_harbor_terminus2.sh

See scripts/miles/README.md for the index, rollout paths and shared code.

Docs

  • Configuration β€” what to change (model, dataset, runtime) and what to leave alone
  • Dataset β€” download and register datasets
  • Evaluation β€” run eval with local or remote Docker
  • Slot Pool Service β€” distribute environments across remote nodes
  • Env Service β€” remote TerminalEnvironment execution on CPU servers
  • Results β€” what each evaluation records and what the fields mean
  • Training β€” AReaL RL training
  • Miles Training β€” Miles RL training examples (DeepSeek-V4, GLM-4.7-Flash, GLM-5.2, Inkling, Qwen3.8) via the Harbor agent server or env_service

Experiments

  • Experiments β€” log of training and evaluation runs

Acknowledgements

The Miles-based RL training pipeline (scripts/miles/, the Harbor agent-server and seta_env session-server wiring, and the sandbox integrations) was built in collaboration with the RadixArk miles team. Thank you for the miles framework and for the support throughout.

Citation

Please cite both the SETA project and the arXiv paper when using this work.

SETA project

@misc{seta,
  author    = {Qijia Shen and Jay Rainton and Aznaur Aliev and Ahmed Awelkair and Boyuan Ma and Zhiqi (Julie) Huang and Yuzhen Mao and Wendong Fan and Philip Torr and Bernard Ghanem and Changran Hu and Urmish Thakker and Guohao Li},
  title     = {{SETA: Scaling Environments for Terminal Agents}},
  year      = {2026},
  month     = jan,
  url       = {https://github.com/camel-ai/seta},
  note      = {Blog: \url{https://eigent-ai.notion.site/SETA-Scaling-Environments-for-Terminal-Agents-2d2511c70ba280a9b7c0fe3e7f1b6ab8}}
}

SETA paper

@misc{shen2026seta,
  author        = {Qijia Shen and Zhiqi Huang and Vamsidhar Kamanuru and Aznaur Aliev and Jay Rainton and Ahmed Awelkair and Zhichen Zeng and Jiajun Li and Shi Dong and Yueming Yuan and Boyuan Ma and Qizheng Zhang and Jiwei Fu and Yuzhen Mao and Wendong Fan and Ping Nie and Philip Torr and Bernard Ghanem and Changran Hu and Jonathan Lingjie Li and Urmish Thakker and Guohao Li},
  title         = {{SETA: Scaling Environments for Terminal Agents}},
  year          = {2026},
  month         = jul,
  eprint        = {2607.10891},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2607.10891}
}

About

πŸ’» SETA: Scaling Environments for Terminal Agents

Resources

Stars

158 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages