From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
cs.LG, cs.AI, cs.CL
Submitted: 2025-05-28
Updated: 2026-08-27
Code: https://github.com/hkust-nlp/RL-Verifier-Robustness
Terminology
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- GPT-4o System Card
- OpenAI o1 System Card
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
- Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Qwen2.5 Technical Report
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks