Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
cs.LG, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/Dodojordi/SP3O
License: http://creativecommons.org/licenses/by/4.0/
The gist: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates.
Terminology
Abstract
In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empirical analyses relate Value Flattening to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states with similar gradients. Motivated by these findings, we introduce SParse Proximal Policy Optimization (SP cubed O), which applies the value loss to only a few well-separated states in each response to mitigate both effects. Experiments on Qwen3-Base show that SP cubed O with only three states supervised per response can mitigate Value Flattening and consistently improve the learned policy across model sizes and evaluation suites. Together, our results identify Value Flattening as an important yet overlooked failure mode of critic learning in standard PPO and show that a simple sparse supervision strategy can mitigate it.
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- OpenAI Gym
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
- P1: Mastering Physics Olympiads with Reinforcement Learning
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- RREDCoT: Segment-Level Reward Redistribution for Reasoning Models
- VIMPO: Value-Implicit Policy Optimization for LLMs
- VinePPO: Refining Credit Assignment in RL Training of LLMs
- Kimi K3: Open Frontier Intelligence
- Solving Quantitative Reasoning Problems with Language Models
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
- CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks