Skip to content

Repository files navigation

Agent Reward GRPO

A minimal end-to-end implementation of the RULER + GRPO workflow: generate trajectories, score them with an LLM-as-judge, normalize group-relative rewards, update a learnable policy, and execute on held-out scenarios. See ARCHITECTURE.md for design details.

Quick Start

Requires an LLM judge. Copy .env.example to .env and fill in your credentials, then run:

Azure OpenAI (recommended — uses az login, no API key needed):

az login
uv run python -m agent_reward_grpo demo --judge azure-openai-entra

OpenAI API:

uv run python -m agent_reward_grpo demo --judge openai --model <your-model>

Outputs written to samples/demo/:

  • data/scenarios.jsonl — train/eval scenarios
  • training/training_metrics.json — per-step rewards and probabilities
  • training/scored_trajectory_groups.json — RULER-scored trajectories
  • policy/learned_policy.json — learned logits and probabilities
  • execution/execution_report.md — held-out evaluation report

Reusing a Learned Policy

Train once, reuse without retraining:

# Train
uv run python -m agent_reward_grpo demo --judge azure-openai-entra --steps 8

# Execute with learned policy (greedy by default)
uv run python -m agent_reward_grpo execute --policy-path samples/demo/policy/learned_policy.json

# Execute sampling from the learned distribution (flexible behavior)
uv run python -m agent_reward_grpo execute --no-greedy --split all

Use --split train, --split eval, or --split all to choose which scenarios to run.

About

Reward-driven agent tuning with GRPO to build adaptive, learnable agents

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages