Bellman Policy Optimization
cs.LG, cs.CL, math.OC
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs).
Terminology
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective. The reformulation avoids estimating state values at intermediate states. We prove that it has the same unique optimal solution as the original PMD objective. We derive the practical BPO loss by approximating this objective. Its mismatch-correction weight is a smoothed ratio of complementary token probabilities. Experiments on mathematical reasoning benchmarks demonstrate the effectiveness of BPO.
Sources
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen3 Technical Report
- Phi-4-reasoning Technical Report
- Group Sequence Policy Optimization
- Rethinking the Trust Region in LLM Reinforcement Learning
- FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
- Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
- VIMPO: Value-Implicit Policy Optimization for LLMs
- Understanding R1-Zero-Like Training: A Critical Perspective
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Proximal Policy Optimization Algorithms
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
- Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
- Defeating the Training-Inference Mismatch via FP16
- Rethinking the Divergence Regularization in LLM RL
- Mirror Descent Policy Optimization
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks