The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment
cs.LG
Submitted: 2026-06-26
Updated: 2026-09-26
Code: https://github.com/openai/prm800k
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- Process Reinforcement through Implicit Rewards
- The Llama 3 Herd of Models
- Qwen2.5-Coder Technical Report
- Mistral 7B
- From Correlation to Causation: Max-Pooling-Based Multi-Instance Learning Leads to More Robust Whole Slide Image Classification
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Learning and Interpreting Multi-Multi-Instance Learning Networks
- Solving math word problems with process- and outcome-based feedback
- Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning
- OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
- Qwen3 Technical Report
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
- Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image Classification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks