Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Your Language Model is Its Own Critic".
Jane: Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We've been diving into the POISE paper, "Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States," and it seems the main thrust is that internal representations are a powerful, practical tool for RL instead of just being diagnostic tools.
Jane: I agree, Tom; the paper really argues that by using a lightweight probe trained online on internal signals, we can estimate the expected verifiable reward without needing an LLM-scale critic.
Lu: The authors show this approach provides a compute-efficient path toward stable and scalable reinforcement learning with verifiable rewards for large reasoning models.
Meng: So, the implication here is that we can achieve better reward variance reduction while using much less computational power during the policy update phase than methods requiring a full critic model.
Lalam: For us at the startup, this means we can potentially train more capable reasoning systems on our existing infrastructure with reduced training costs.
Tom: Exactly; they also show that this method avoids the extra sampling needed to identify and discard prompt groups that have zero advantage by providing a lightweight continuous baseline for each rollout.
Jane: And it seems the overall implication is that we can achieve performance comparable to existing state-of-the-art algorithms on mathematical reasoning benchmarks while requiring less compute.
Lu: This suggests that the way we structure our RL training—how we estimate value functions—is a key area where these internal signals can be used to create more robust and scalable systems.
Meng: I'm curious about what the authors pointed out as a specific limitation in their work; what’s the method not able to do or where it stops working?.
Lalam: The paper notes that this method relies on training a lightweight probe, which means its performance is tied to how well that initial probe learns from the internal signals.
Tom: That's a fair point; the success of POISE hinges on the quality and accuracy of that initial value estimator, as it doesn't magically solve every RL problem.
Jane: So, in simple terms, "Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States" suggests we can use the model’s own internal workings to build more stable and resource-efficient reinforcement learning for large reasoning models.
Conclusion: Tom: So, to wrap up our discussion on "Your Language Model is Its Own Critic," we've seen how this new method uses an AI's internal workings to create a more stable way to train models through reinforcement learning.
Jane: That’s right, Tom; the core idea is moving away from needing some massive external critic model and instead using signals already inside the model itself to estimate what a good outcome looks like.
Lu: The authors are really smart for turning something usually treated as just a diagnostic tool into an active signal for policy improvement in this way.
Meng: From my side, it’s interesting because it sounds like we can get better performance without having to scale up the entire training setup with a huge critic component, which saves us time and resources on the compute side.
Lalam: For me, this has huge cultural implications; if we can make reasoning training more stable and efficient this way, it means our AI culture evolves in a much safer and more predictable manner.
Tom: Exactly; the title itself really gets to the heart of it—the model becomes its own critic, which is a neat concept for how we think about training these large systems.
Jane: It’s simple to grasp, Tom; basically, instead of waiting for an outside expert to tell us if our AI is doing well, we train a small component inside the AI that learns to predict the reward based on what it's already thinking.
Lu: I think what’s most exciting is how this connects the internal state features—the prompt and trajectory signals—directly into a value estimation framework that's tailored specifically for language models.
Meng: That connection is key, because it allows us to use these internal representations in a way that directly guides the policy optimization steps, making the learning process much more direct.
Lalam: And this efficiency means we can explore more complex reasoning tasks without hitting those frustrating plateaus where training just stalls because the value estimation isn't accurate enough.
Tom: Indeed; so, the authors of this paper are showing us a way to make reinforcement learning for large language models less computationally intensive while keeping it grounded in what the model actually learns internally.
Jane: It’s about making that internal knowledge actionable for policy updates, which is a big step forward in how we design these complex AI systems.
Lu: And this opens up so many creative avenues; imagine using those internal states not just for value estimation but for novel ways to structure the entire reasoning process itself.
Meng: I’m thinking about how this might affect deployment; if training is more stable, the models that come out are probably going to be more reliable in real-world applications.
Lalam: That reliability translates directly into trust, and for our team, it means we can focus our energy on building richer capabilities rather than fighting unstable learning dynamics.
Tom: So, this paper really lays out a pathway where the internal signals of the AI become a direct and useful source of guidance for reinforcement learning.
Graduate School of Data Science Seoul National University
cs.LG, cs.AI, cs.CL
Submitted: 2026-05-08
Updated: 2026-10-03
Comments: Accepted to NeurIPS 2026; Project Page: https://holi-lab.github.io/POISE/
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions, a
Key concepts
- Internal Representations
- These are the hidden states computed during policy training that capture how the language model understands prompts and reasoning steps. Instead of just looking at the final output, these internal signals are used as rich optimization signals for reinforcement learning.
- Lightweight Probe
- This is a small neural network trained online to predict the expected reward from internal states. It takes features derived from prompt and reasoning hidden states, along with token entropy statistics, to estimate the value function Vpi(x). This probe acts as an efficient critic without needing a full LLM-scale critic.
- Cross-Rollout Construction
- This technique involves sampling two independent rollouts for each prompt. The baseline for one rollout is predicted using the internal signals from the other rollout. This ensures that the advantage calculation remains unbiased and conditional on the sampled data, preserving gradient accuracy during policy updates.
Terminology
Summary
Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions, a concept introduced in POISE. The core finding is that a lightweight probe trained online on these internal signals can predict the expected verifiable reward from hidden states, enabling a baseline estimation method that matches state-of-the-art algorithms while requiring less compute and offering more stable learning dynamics.
The gist
Internal representations of reasoning models can move beyond their conventional use as diagnostic tools for reasoning behavior and serve as practical optimization signals for reinforcement learning.
How it works: Value Function Estimation from Policy Model Internal States
The paper introduces POISE, an RL algorithm that turns the model’s internal states into a value model by training a lightweight probe to predict the prompt-level value, defined as the expected verifier reward, denoted as Vπ(x) = E[y π(·x)[R(x, y)]. This is achieved by training a lightweight probe
on internal signals collected at two levels:
-
A prompt-level feature: extracted from hidden states at the final prompt tokens before generation begins, capturing how the model represents the prompt and its anticipated difficulty.
-
A trajectory-level feature: comprising hidden states taken when the model’s reasoning ends together with token-level entropy statistics.
The probe input is constructed as a combination of these signals: the prompt-state feature h(i)θ,p = Avgt∈P H(i)θ,t
and the reasoning-state feature h(i)θ,r = Avgt∈R(i)H(i)θ,t,
along with token-level entropy statistics u(i)θ.
The probe is trained to minimize the loss Lvalue = Exi gf (ϕ(i)θ) − Vb−i(x), where the target Vb−i(x) is derived from a leave-one-out Monte Carlo target, ensuring conditional independence of the target from the features used by the probe.
How it works: Policy Optimization with Cross-Rollout Baselines
The POISE algorithm integrates this value estimator into policy optimization by using a cross-rollout construction
to preserve gradient unbiasedness. For each prompt, two independent rollouts, y(1) and y(2), are sampled from the current policy πθold. The baseline for each rollout is predicted from the internal signals of the other rollout: b(1)(x) = gfϕ(2))
and b(2)(x) = gfϕ(1)).
This construction ensures that the baseline used to update y(i) depends only on the independently sampled rollout y(j), j ≠ i, satisfying the conditional-independence condition in Eq. (6).
The resulting advantages are calculated as A(i)(x) = R(x, y(i)) − b(i)(x)."
The policy is then updated using a PPO-style clipped surrogate objective:
Lθ = Ex∼D, y(1),y(2)∼πθ"
1/2 X 2 i=1 y(i) y(i) X t=1 min n n r(i) t θ A(i)(x), clip r (i) t θ, 1 − ε, 1 + ε A(i)(x)"
How it works: Online Estimator Training with a Trajectory Buffer
The value estimator gf is trained jointly with the policy on a sliding buffer of recent trajectories.
At each step, for each prompt x with two independent rollouts (y(1), y(2)), the value estimator examples are constructed using the paired rollout reward Re(i) = R(x, y(j)), j ≠ i
and corresponding internal features. The update minimizes a regression loss over this buffer and new examples: Lvalue(f) = Ex,i gf (ϕ(i)θ) − re2.
This online training allows the estimator to adapt to policy changes while maintaining negligible overhead.
Key Advantages and Performance
POISE offers several concrete advantages over existing methods:
-
It avoids an LLM-scale critic, using a
lightweight value estimator rather than an LLM-scale critic.
-
Compared to GRPO, it requires only
a pair of rollouts rather than a large group,
allowing the saved budget to be redirected tomore distinct prompts per batch.
-
It avoids the extra sampling needed to identify and discard degenerate zero-advantage prompt groups because it provides a
lightweight continuous baseline for each rollout.
Empirically, POISE matches DAPO on mathematical reasoning benchmarks while requiring less compute.
Improvements for AI systems
Here are specific improvements for AI systems based on the POISE (Policy Optimization with Internal State Value Estimation) framework:
-
Incorporate a lightweight, online value estimator directly into Reinforcement Learning (RL) policy optimization loops, eliminating the need for computationally expensive LLM-scale critics or large groups of rollouts per prompt.
-
Enable more stable and efficient RLVR (Reinforcement Learning with Verifiable Rewards) for Large Reasoning Models by replacing traditional variance reduction baselines (like PPO's LLM-scale critic or GRPO's multiple rollouts) with a baseline derived from the policy model’s own internal signals.
-
Improve prompt diversity and training stability by allowing the use of fewer rollouts per prompt, which can be reallocated to explore more distinct prompts within a fixed compute budget, thereby reducing gradient variance.
-
Develop an AI system capable of performing complex mathematical reasoning tasks (like those in AMC, AIME, or BRUMO) with higher accuracy and stability by leveraging the
prompt-level value
estimated from hidden states before generation begins. -
Create a general-purpose RLVR mechanism that can be applied across diverse verifiable domains beyond math reasoning, such as tool-calling dialogues (ToolDial), instruction following (IF-RLVR), and coding tasks (AceCoder).
-
Enhance the robustness of reasoning models by utilizing trajectory-level hidden states, combined with token-entropy statistics, to provide richer value signals for the internal state probe.
-
Develop an AI that can dynamically track nonstationary expected rewards during continuous RL training, ensuring the learned baseline remains calibrated to the evolving policy distribution without requiring frequent retraining of a separate critic model.
-
Improve inference efficiency by allowing models to utilize their own internal representations (hidden states) as cheap, pre-computed value estimates during the forward pass, potentially leading to faster or more reliable decision-making processes during generation.
Abstract
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimation), a reinforcement learning algorithm that turns the model's internal states into a value model. A lightweight probe reads the signals already computed during the forward pass to predict the baseline, and is trained online alongside the policy. To preserve gradient unbiasedness, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. On Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines while achieving more stable training. Moreover, the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. By leveraging the model's internal representations, POISE enables stable policy optimization.
Sources
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- Prompt Curriculum Learning for Efficient LLM Post-Training
- OpenAI o1 System Card
- Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
- Single-stream Policy Optimization
- Qwen3 Technical Report
- Demystifying Long Chain-of-Thought Reasoning in LLMs
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- $V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks