Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
summary
The gist
Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions, a
In short
The paper introduces POISE, an RL method that uses a lightweight probe trained online on a model's internal states to estimate value functions. This allows the system to predict expected rewards from hidden states, creating a stable baseline estimation method. This approach matches state-of-the-art performance while reducing computational requirements and improving learning stability.
Key concepts
- Internal Representations
- These are the hidden states computed during policy training that capture how the language model understands prompts and reasoning steps. Instead of just looking at the final output, these internal signals are used as rich optimization signals for reinforcement learning.
- Lightweight Probe
- This is a small neural network trained online to predict the expected reward from internal states. It takes features derived from prompt and reasoning hidden states, along with token entropy statistics, to estimate the value function Vpi(x). This probe acts as an efficient critic without needing a full LLM-scale critic.
- Cross-Rollout Construction
- This technique involves sampling two independent rollouts for each prompt. The baseline for one rollout is predicted using the internal signals from the other rollout. This ensures that the advantage calculation remains unbiased and conditional on the sampled data, preserving gradient accuracy during policy updates.
Terminology used across episodes
This episode discusses
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States · Paper Radio
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- Prompt Curriculum Learning for Efficient LLM Post-Training
- OpenAI o1 System Card
- Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
- Single-stream Policy Optimization
- Qwen3 Technical Report
- Demystifying Long Chain-of-Thought Reasoning in LLMs
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- V 0.5: Generalist Value Model as a Prior for Sparse RL Rollouts
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- Representation Engineering: A Top-Down Approach to AI Transparency
The paper
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States · Read on arXiv
Graduate School of Data Science Seoul National University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Your Language Model is Its Own Critic".
Jane: Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We've been diving into the POISE paper, "Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States," and it seems the main thrust is that internal representations are a powerful, practical tool for RL instead of just being diagnostic tools.
Jane: I agree, Tom; the paper really argues that by using a lightweight probe trained online on internal signals, we can estimate the expected verifiable reward without needing an LLM-scale critic.
Lu: The authors show this approach provides a compute-efficient path toward stable and scalable reinforcement learning with verifiable rewards for large reasoning models.
Meng: So, the implication here is that we can achieve better reward variance reduction while using much less computational power during the policy update phase than methods requiring a full critic model.
Lalam: For us at the startup, this means we can potentially train more capable reasoning systems on our existing infrastructure with reduced training costs.
Tom: Exactly; they also show that this method avoids the extra sampling needed to identify and discard prompt groups that have zero advantage by providing a lightweight continuous baseline for each rollout.
Jane: And it seems the overall implication is that we can achieve performance comparable to existing state-of-the-art algorithms on mathematical reasoning benchmarks while requiring less compute.
Lu: This suggests that the way we structure our RL training—how we estimate value functions—is a key area where these internal signals can be used to create more robust and scalable systems.
Meng: I'm curious about what the authors pointed out as a specific limitation in their work; what’s the method not able to do or where it stops working?.
Lalam: The paper notes that this method relies on training a lightweight probe, which means its performance is tied to how well that initial probe learns from the internal signals.
Tom: That's a fair point; the success of POISE hinges on the quality and accuracy of that initial value estimator, as it doesn't magically solve every RL problem.
Jane: So, in simple terms, "Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States" suggests we can use the model’s own internal workings to build more stable and resource-efficient reinforcement learning for large reasoning models.
Conclusion: Tom: So, to wrap up our discussion on "Your Language Model is Its Own Critic," we've seen how this new method uses an AI's internal workings to create a more stable way to train models through reinforcement learning.
Jane: That’s right, Tom; the core idea is moving away from needing some massive external critic model and instead using signals already inside the model itself to estimate what a good outcome looks like.
Lu: The authors are really smart for turning something usually treated as just a diagnostic tool into an active signal for policy improvement in this way.
Meng: From my side, it’s interesting because it sounds like we can get better performance without having to scale up the entire training setup with a huge critic component, which saves us time and resources on the compute side.
Lalam: For me, this has huge cultural implications; if we can make reasoning training more stable and efficient this way, it means our AI culture evolves in a much safer and more predictable manner.
Tom: Exactly; the title itself really gets to the heart of it—the model becomes its own critic, which is a neat concept for how we think about training these large systems.
Jane: It’s simple to grasp, Tom; basically, instead of waiting for an outside expert to tell us if our AI is doing well, we train a small component inside the AI that learns to predict the reward based on what it's already thinking.
Lu: I think what’s most exciting is how this connects the internal state features—the prompt and trajectory signals—directly into a value estimation framework that's tailored specifically for language models.
Meng: That connection is key, because it allows us to use these internal representations in a way that directly guides the policy optimization steps, making the learning process much more direct.
Lalam: And this efficiency means we can explore more complex reasoning tasks without hitting those frustrating plateaus where training just stalls because the value estimation isn't accurate enough.
Tom: Indeed; so, the authors of this paper are showing us a way to make reinforcement learning for large language models less computationally intensive while keeping it grounded in what the model actually learns internally.
Jane: It’s about making that internal knowledge actionable for policy updates, which is a big step forward in how we design these complex AI systems.
Lu: And this opens up so many creative avenues; imagine using those internal states not just for value estimation but for novel ways to structure the entire reasoning process itself.
Meng: I’m thinking about how this might affect deployment; if training is more stable, the models that come out are probably going to be more reliable in real-world applications.
Lalam: That reliability translates directly into trust, and for our team, it means we can focus our energy on building richer capabilities rather than fighting unstable learning dynamics.
Tom: So, this paper really lays out a pathway where the internal signals of the AI become a direct and useful source of guidance for reinforcement learning.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck