A minimal end-to-end implementation of the RULER + GRPO workflow: generate trajectories, score them with an LLM-as-judge, normalize group-relative rewards, update a learnable policy, and execute on held-out scenarios. See ARCHITECTURE.md for design details.
Requires an LLM judge. Copy .env.example to .env and fill in your credentials, then run:
Azure OpenAI (recommended — uses az login, no API key needed):
az login
uv run python -m agent_reward_grpo demo --judge azure-openai-entraOpenAI API:
uv run python -m agent_reward_grpo demo --judge openai --model <your-model>Outputs written to samples/demo/:
data/scenarios.jsonl— train/eval scenariostraining/training_metrics.json— per-step rewards and probabilitiestraining/scored_trajectory_groups.json— RULER-scored trajectoriespolicy/learned_policy.json— learned logits and probabilitiesexecution/execution_report.md— held-out evaluation report
Train once, reuse without retraining:
# Train
uv run python -m agent_reward_grpo demo --judge azure-openai-entra --steps 8
# Execute with learned policy (greedy by default)
uv run python -m agent_reward_grpo execute --policy-path samples/demo/policy/learned_policy.json
# Execute sampling from the learned distribution (flexible behavior)
uv run python -m agent_reward_grpo execute --no-greedy --split allUse --split train, --split eval, or --split all to choose which scenarios to run.