ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth
cs.LG, cs.AI, cs.CL
Submitted: 2025-12-02
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale
Terminology
Abstract
Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervision. We propose ScalePRM, which scales verification compute as an alternative: given a problem and a candidate solution, we generate multiple independent verifications of each reasoning step and aggregate their judgments to produce synthetic step-level labels without ground truth. We explore two representative inference-time scaling strategies, parallel scaling through self-consistency and sequential scaling through meta-critique, and train generative PRMs on the resulting synthetic data. On ProcessBench, a benchmark for identifying erroneous steps in mathematical reasoning, PRMs trained on step-level self-consistency data achieve 67.5 F1, surpassing reference-guided training with ground-truth access (66.4 F1) and GPT-4o as a critic (61.9 F1). When deployed as reward signals in RL training with Qwen2.5-Math-7B, our best PRM achieves 47.4% average accuracy across six mathematical reasoning benchmarks, outperforming ground-truth-based RLVR (43.9%). We also identify and address reward exploitation patterns unique to generative PRM-based RL. Our results demonstrate that scaling verification compute is a viable alternative to ground-truth supervision for training process reward models.
Sources
- Measuring Progress on Scalable Oversight for Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Training Verifiers to Solve Math Word Problems
- Process Reinforcement through Implicit Rewards
- Skywork Open Reasoner 1 Technical Report
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- GPT-4o System Card
- OpenAI o1 System Card
- LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- Process Reward Models That Think
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Decoupled Weight Decay Regularization
- Categorizing Variants of Goodhart's Law
- LLM Critics Help Catch LLM Bugs
- Self-critiquing models for assisting human evaluators
- Spurious Rewards: Rethinking Training Signals in RLVR
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks