Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".
-
Updated
Feb 19, 2025 - Python
Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".
[ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.
This repository contains the official implementation of Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision.
[KDD 2025] Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
[NeurIPS 2025 Spotlight] LLM post-training suite — featuring ReasonFlux, ReasonFlux-PRM, and ReasonFlux-Coder.
[AAAI 2026] Official codebase for "GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning".
Interface for DeepSeek-R1 chain-of-thought model — exposes the internal reasoning trace, supports API and local Ollama backends, with streaming output and temperature control.
Live proof of arXiv:2603.17815 — O(N) confirmed R²=0.952, 1,984 API calls
This repository includes code and materials for the paper "Efficient PRM Training Data Synthesis via Formal Verification" (ACL 2026 Findings).
Step-level preference data construction toolkit for code-agent process reward models
Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).
[ACL 2026 SAC Highlight Award] Can We Predict Before Executing Machine Learning Agents?
DPO-trained decision layer for coding agents — raises pass@1 from 0.419 → 0.639 on frozen 14B generator (p<0.000001)
Multi-Agent System Process Reward Model (MASPRM): a lightweight process reward model guiding multi-agent systems at search time.
Code for 'Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs' — a rigorously established inference-time mechanism null (3 LMs, 6 benchmarks, 2 PRMs, N=8–256, pre-registered TOST)
A comprehensive collection of process reward models.
Official code for CulturePRM: A Process Reward Model for Mitigating Cultural Overriding in Cultural Reasoning (EMNLP 2026 Main)
Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.
To associate your repository with the process-reward-model topic, visit your repo's landing page and select "manage topics."