Verify to Amplify: Improving Reasoning via Learned Chain-of-Thought Verification
cs.LG
Submitted: 2026-03-03
Updated: 2026-09-09
Comments: The abstract has been abridged due to arXiv length constraints
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Process Reinforcement through Implicit Rewards
- A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
- Towards Autonomous Mathematics Research
- Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift
- Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
- AI safety via debate
- Process Reward Models That Think
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Let's reward step by step: Step-Level reward model as the Navigators for Reasoning
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
- Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
- Solving math word problems with process- and outcome-based feedback
- When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- The Lessons of Developing Process Reward Models in Mathematical Reasoning
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks