Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature Annealing
stat.ML, cs.LG
Submitted: 2025-12-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fitted Q-iteration (FQI) is a standard regression-based method for optimal control in offline reinforcement learning, but its stability under function approximation often relies on Bellman
Terminology
Abstract
Fitted Q-iteration (FQI) is a standard regression-based method for optimal control in offline reinforcement learning, but its stability under function approximation often relies on Bellman completeness, which requires Bellman images of the fitted class to remain in the class. We study Kullback--Leibler (KL)-regularized, or soft, FQI relative to a fixed reference policy without this assumption. Our key insight is that soft control locally inherits the contraction of policy evaluation in a discounted-occupancy norm. At the soft-optimal fixed point, the linearization of the soft Bellman operator is exactly the Bellman operator for the soft-optimal policy, which contracts in its discounted-occupancy norm; projection in the same norm preserves this contraction. Standard soft FQI instead projects under the offline state-action distribution and need not preserve this property. Motivated by this observation, we propose occupancy-reweighted soft FQI, which retains standard Bellman targets and least-squares updates while reweighting regressions by discounted-occupancy ratios induced by the current soft policy. Under Q-function realizability and local regularity, we establish local contraction and finite-sample convergence with estimated ratios, without Bellman completeness. We then use temperature annealing to convert the local result into global convergence from arbitrary initialization: sufficiently high temperature provides a globally contractive starting regime, while gradual cooling connects successive local contraction regions to any prescribed positive target temperature. Under an action-gap margin condition, switching at a fixed positive temperature to hard FQI with refreshed occupancy weights also yields population and finite-sample convergence to the unregularized optimum.
Sources
- A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting
- Harnessing Density Ratios for Online Reinforcement Learning
- Sequential causal inference in a single world of connected units
- Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
- Emphatic Temporal-Difference Learning
- Off-Policy Evaluation in Markov Decision Processes under Weak Distributional Overlap
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- A unified view of entropy-regularized Markov decision processes
- Source-Condition Analysis of Kernel Adversarial Estimators
- Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
- A Researcher's Guide to Empirical Risk Minimization
- Nonparametric Instrumental Variable Inference with Many Weak Instruments
- Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning
- The Role of Coverage in Online Reinforcement Learning
- GenDICE: Generalized Offline Estimation of Stationary Values
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey