Stepwise Intrinsic Rewards for Reasoning in Large Language Models
cs.AI, cs.CL
Submitted: 2026-02-01
Updated: 2026-09-25
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Bootstrapping Language Models with DPO Implicit Rewards
- G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
- Reinforced Self-Training (ReST) for Language Modeling
- Training Verifiers to Solve Math Word Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- OpenAI o1 System Card
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Are NLP Models really able to Solve Simple Math Word Problems?
- Kimi K2: Open Agentic Intelligence
- Qwen2.5 Technical Report
- The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Trust-Region Adaptive Policy Optimization
- Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
- STaR: Bootstrapping Reasoning With Reasoning
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection