Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Hidden States as Value Gradients".
Jane: A key capability of intelligent agents is to act effectively under incomplete state observations,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We've established the core concept linking hidden states to the co-state structure, so now let’s talk about who put this research out there and what the title itself tells us about this work. Jane, can you give us a quick rundown of the paper's title and authors?
Jane: Certainly, Tom. The paper is titled "Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies," and it was written by David Leeftink, Max Hinne, and Marcel van Gerven from the Donders Institute for Brain, Cognition and Behaviour at Radboud University.
Lu: Those authors are deep in the realm of machine learning and neural computing, so you expect a lot of theoretical rigor here; it’s clear they are looking to bridge those two worlds.
Meng: I wonder if their background in brain cognition helps them see this connection between recurrent memory and control theory more clearly than someone purely focused on optimization algorithms would.
Lalam: It’s interesting to see that the people working on these deep learning models are drawing connections back to classical optimal control theory, which feels like a big step for integrating different AI disciplines.
Tom: It really is; it shows that the problems of effective action under incomplete observation aren't just solved by deeper networks, but by realizing specific mathematical structures within their memory mechanisms.
The paper's summary: Jane: Now, moving onto the actual summary of "Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies," the authors explain that recurrent policies address incomplete observation by compressing histories into a hidden state, but they demonstrate this state plays the role of the co-state in Pontryagin’s minimum principle.
Lu: They show that along optimal trajectories, this co-state is precisely the value gradient, which evolves affinely in time according to Pontryagin's principle: = -grad x q - (grad xx) lambda (<ref:2605.05373#pg1>).
Tom: That’s a lot of math to unpack, but the main point is that they formalize this correspondence by defining a class of policies called costate policies, which are defined by having an affine recurrence in the hidden state and performing Hamiltonian minimization in their readout.
Jane: So, it’s about showing that these structural properties automatically force the policy to realize this optimal control structure, meaning we don't have to manually design every single thing for optimality.
Meng: That sounds powerful because if the architecture itself is designed correctly, it should inherently capture a lot of the necessary information about the environment's sensitivity without needing massive amounts of auxiliary data.
Lalam: It’s like discovering an inherent organizational principle in how memory works that guides it toward achieving better outcomes, which is something we can really leverage for future AI development.
The paper's improvements: Tom: So, if the paper shows this correspondence and introduces costate policies, what are the specific improvements they propose for actual training and architecture? Jane, can you summarize the key enhancements suggested by the authors?
Jane: The primary improvement is introducing a co-state loss for actor-critic training where the critic’s gradient serves as a target signal for shaping the actor’s hidden state. This is formalized through an alignment objective, Lco-state(θ, W) = Et h W⊤ ht − sg λˆt 2i (<ref:2605.05373#pg0>).
Lu: That loss function is clever because it aligns the two through a linear embedding, making it mathematically tractable and allowing us to supervise the hidden state dynamics directly with the critic's signal.
Meng: From an engineering viewpoint, this means we can potentially train agents to track optimality structures directly, which could stabilize training significantly compared to just relying on standard reward maximization alone.
Tom: And they also showed that empirically, this objective performs well on challenging locomotion tasks like the H1 and Berkeley humanoids when compared against just reward maximization.
Jane: Plus, they designed a readout layer that minimizes the control-Hamiltonian, which means the action selection isn't arbitrary; it’s mathematically determined by a switching function derived from the hidden state and observation.
Conclusion: Tom: We’re coming to the end of our discussion on "Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies." So, let's summarize the main implications for us before we wrap things up. Jane, can you give us your final thoughts on what this paper means for our field?
Jane: What this paper suggests is that the hidden state acts as a Pontryagin co-state, providing a control-theoretic account of what recurrent policies compute in continuous control problems and offering a mechanism to shape those states toward optimality via the co-state loss.
Lu: It opens up the door to designing recurrent architectures where memory has an explicit structure tied to sensitivity, which is a powerful theoretical direction for building more capable agents.
Meng: I think the practical implication is that we can expect training methods that directly supervise these structural properties of our memory, leading to better performance on real-world robotic control tasks and potentially more robust systems.
Lalam: For us, it’s about developing AI that possesses an internal representation of what it means to be optimal in action, which could make our overall system behavior feel much more coherent and predictable.
Tom: It’s been a fascinating look at how we can connect the minimum principle to recurrent memory through this paper on "Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies." Thanks for hanging out with us, everyone.
Department of Machine Learning and Neural Computing · Donders Institute for Brain, Cognition and Behaviour · Radboud University
cs.LG
Submitted: 2026-05-06
Updated: 2026-09-28
Comments: 19 pages, 8 figures
Code: https://github.com/DavidLeeftink/anonymousfornow
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: A key capability of intelligent agents is to act effectively under incomplete state observations, and this work shows that recurrent policies implicitly realize a control-theoretic structure where
Key concepts
- Pontryagin’s Minimum Principle (PMP)
- A fundamental principle in optimal control that helps determine the best path for a system. It involves defining an adjoint variable, or co-state, which measures how sensitive the final outcome is to changes in the current state. The optimal control action is chosen to minimize a specific Hamiltonian function derived from this principle.
- Co-state (λ)
- In optimal control, the co-state is a crucial mathematical tool that tracks the sensitivity of the optimal value function with respect to state variables. The paper argues that in continuous control, this co-state corresponds to the gradient of the value function ($ abla_x V(x)$), which dictates how much future rewards change based on current state values.
- Costate Policy (CP)
- A specific type of recurrent policy designed to mimic optimal control behavior. A CP has two properties: its recurrence must be affine in the hidden state, and the readout layer must perform Hamiltonian minimization. This structure allows the hidden state to realize the optimal co-state dynamics required for an optimal trajectory.
- Co-state Loss
- A training objective used in actor-critic methods where a loss function is defined to align the actor's hidden state with a target derived from the critic's gradient. This loss explicitly forces the hidden state to track the co-state, providing a mechanism for supervising optimality beyond simple reward maximization.
Terminology
Summary
A key capability of intelligent agents is to act effectively under incomplete state observations, and this work shows that recurrent policies implicitly realize a control-theoretic structure where their hidden state tracks the gradient of the value function, playing the role of a co-state in Pontryagin’s minimum principle. This correspondence provides a mechanism for shaping hidden states toward optimality through a co-state loss, connecting recurrent memory to the optimality structure of continuous control problems.
The core theoretical framework
The paper formalizes this relationship by drawing a link between recurrent policies and the Pontryagin minimum principle (PMP). In optimal control theory, the co-state measures the sensitivity of the optimal value function to the current state,
and it evolves according to affine dynamics. The authors argue that in continuous control, this co-state coincides with the gradient of the value function, ∇xV (x),
which appears in Pontryagin’s principle. They introduce a class of policies called costate policies (CPs), defined by two structural properties: (1) that the recurrence is affine in the hidden state,
and (2) that the readout performs Hamiltonian minimization.
The co-state policy structure
A co-state policy (CP) is formally defined as a recurrent policy where the update rule is given by equation (3.3):
h˙ = −bθ(y) − Fθ(y, u)h, u = arg min u∈U ⊤Gθ(y)h.
The paper demonstrates that for certain architectures like structured state space models (SSMs) and linear RNNs, this structure is realized. Specifically, Proposition 3.3 shows that if the network operators satisfy certain conditions (Eq. 3.4), the hidden state realizes the optimal co-state dynamics h(t) = T λ⋆(t)
for all time along an optimal trajectory, and emits the optimal action u⋆(t).
Supervising optimality via loss functions
The correspondence allows for a co-state loss for actor-critic training,
where the critic’s gradient serves as a target for the actor’s hidden state.
This is formalized by aligning the two through a linear embedding:
Lco-state(θ, W) = Et h W⊤ ht − sg λˆt 2i,
where λˆt is constructed from the critic's gradient. Empirically, they find that this objective lifts nearly all of them above this ceiling
when compared to reward maximization alone on challenging locomotion tasks like the H1 and Berkeley humanoids.
Empirical findings on hidden state encoding
Experiments test whether hidden states encode co-state information beyond what is linearly decodable from the environment state alone. They distinguish between two hypotheses: either h strictly estimates the unobserved state, or it tracks the co-state λ explicitly.
On non-linear stabilization tasks, they find that most hidden states do not encode co-state structure beyond what is linearly decodable from the state,
but that the co-state objective lifts nearly all of them above this ceiling.
Furthermore, they observe that Improved co-state alignment is accompanied throughout by improved state tracking,
consistent with the analysis in Section 3.3.
Readout layer design
The readout mechanism is shown to be determined by minimizing the control-Hamiltonian, which depends on the co-state and is captured by a switching function σ(t):= −Gθ(y)h(t). This allows for the structural design of readouts tailored to different control objectives:
-
For Quadratic Control (Smooth Action): The optimal control law is derived as u⋆ = R−1σ.
-
For Time/State (Bang-Bang): The optimal control is u⋆ = sign(σ)umax if σ > 0, and 0 if σ = 0.
-
For Fuel (Bang-Off-Bang): The optimal control develops a deadzone where it is optimal to coast when σ < 1.
Conclusion
The work connects the minimum principle to recurrent memory, providing a control-theoretic account of what hidden states compute in continuous control and a mechanism for shaping them toward optimality.
They conclude that the hidden state acts as the Pontryagin co-state, while the readout minimizes the control-Hamiltonian.
The results suggest that Directly supervising the hidden state with the value gradient improves both co-state representation and mean return.
The gist
The hidden state of a recurrent policy plays the role of a co-state in Pontryagin’s minimum principle, and this correspondence allows for a co-state loss for actor-critic training that improves policies on challenging locomotion tasks.
How it works
Improvements for AI systems
Here are the specific improvements to AI systems based on this research, categorized by architectural enhancement and training methodology:
) Architectural Enhancement: Implementing Co-state Policies (CPs)
The core improvement is replacing standard recurrent cells (like LSTMs or GRUs) with architectures that explicitly satisfy the Pontryagin Minimum Principle.
- Implement a recurrence relation where the hidden state dynamics are affine in the hidden state, matching the co-state evolution:
[1]: The system should be designed such that the hidden state update is of the form:
ḣ = -bθ(y) - Fθ(y, u)h
where Fθ is a linear transformation (or structured like a Linear SSM/Mamba cell), ensuring the recurrence mirrors the co-state dynamics.
- Implement a readout layer that performs Hamiltonian minimization rather than arbitrary action selection:
[2]: The policy readout should be defined as minimizing the control-Hamiltonian, which structurally corresponds to selecting an optimal control law based on a switching function derived from the hidden state and observation (e.g., using sign functions for bang-bang controls or quadratic forms for smooth controls).
) Training Methodology Enhancement: Co-state Loss Supervision
The primary training improvement is integrating the critic's value gradient directly into the actor's memory update mechanism.
- Introduce a co-state loss term in the actor’s objective function:
[3]: The actor’s hidden state update should be augmented to minimize a loss based on projecting its representation onto the critic’s gradient (the proxy for the value gradient):
h actor target = Lco-state(θ, W) = E[h]W - sg(λ̂t)
where λ̂t is derived from the critic's state/value function, ensuring the actor's memory learns to track optimality structures beyond simple state estimation.
) System Capability Improvements (What the Improved AI Can Do):
By implementing these changes, the improved AI systems will exhibit:
-
Optimal Control in Partially Observable Environments (POMDPs): The agent can perform continuous control tasks (like locomotion or robotics) under noisy, partial observations with significantly higher efficiency and stability. This is because the hidden state is explicitly structured to represent the necessary
sensitivity
of the optimal cost-to-go rather than just a raw belief over states. -
Improved Robustness in Complex Dynamics: The system will be better equipped to handle non-linear dynamics (like those in humanoid locomotion) and contact discontinuities, as the co-state objective specifically targets features that reward maximization alone fails to capture.
-
Direct Mapping of Optimality: The agent’s internal memory structure becomes a direct computational representation of the underlying optimal control problem's sensitivity (the co-state), allowing it to directly solve for near-optimal control laws in real-time, rather than relying on complex, learned mappings from state history alone.
Sources
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Were RNNs All We Needed?
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Efficiently Modeling Long Sequences with Structured State Spaces
- Variational Recurrent Models for Solving Partially Observable Control Tasks
- Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Proximal Policy Optimization Algorithms
- RSL-RL: A Learning Library for Robotics Research
- Simplified State Space Layers for Sequence Modeling
- DeepMind Control Suite
- MuJoCo Playground
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks