Exact-oracle conformance tests for turn-level credit assignment in agentic RL (GRPO, RLOO, GAE, GiGPO; verl, TRL, OpenRLHF).
-
Updated
Sep 17, 2026 - Python
Exact-oracle conformance tests for turn-level credit assignment in agentic RL (GRPO, RLOO, GAE, GiGPO; verl, TRL, OpenRLHF).
Readable 25.7M-parameter GLM-5.3-Flash-style Hybrid-MoE built from scratch in PyTorch. Features hybrid linear/sparse attention, executable-reward RLOO, and recursive self-improvement on verified programs.
MiniCPM5-2B shopping agent, end-to-end: collection → filtering → curriculum SFT → GRPO/RLOO (veRL) → deterministic Final-200 eval. Strict success 3.0% → 76.5%.
RL + human-preference reward model for automatic product lighting in Blender (NN-inherit → 2-step RLOO refine → best-of-16 RM arbitration)
Reproduces Hugging Face's RLOO vs. PPO win-rate comparison on real public checkpoints, using Claude as judge instead of GPT-4. Evaluation only — no training.
To associate your repository with the rloo topic, visit your repo's landing page and select "manage topics."