Lifted Bellman Linear Programming for Offline Reinforcement Learning
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by/4.0/
The gist: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA)
Terminology
Abstract
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint (Q,V) space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along K-step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and horizon. Under deterministic dynamics, this minimizer lies between the best dataset return and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the K-step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it can be trained without target networks or EMA updates. Under deterministic dynamics, the LBLP solution is a stationary point of the detached update under a coefficient condition independent of γ and K, and the inequality constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and is comparable to recent action-chunking methods, while using the fewest parameters and the least peak GPU memory among all compared methods.
Sources
- Learning from Sparse Offline Datasets via Conservative Density Estimation
- IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
- Gaussian Error Linear Units (GELUs)
- Stochastic Primal-Dual Q-Learning
- Adam: A Method for Stochastic Optimization
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Continuous control with deep reinforcement learning
- Convex Q-Learning, Part 1: Deterministic Optimal Control
- Reinforcement Learning via Fenchel-Rockafellar Duality
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- CAQL: Continuous Action Q-Learning
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
- The In-Sample Softmax for Offline Reinforcement Learning
- GenDICE: Generalized Offline Estimation of Stationary Values
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks