Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Lars van der Laan, Nathan Kallus
stat.ML, cs.LG
Submitted: 2026-07-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation.
Terminology
Abstract
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback-Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to approximation error and a statistical error governed by the complexity of the ratio hypothesis class. When full coverage fails, we introduce coverage-stopped FORE, which targets the discounted occupancy accumulated before the first uncovered state-action pair and yields a conservative lower bound on target-policy value for nonnegative rewards. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.
Sources
- A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting
- Source Condition Double Robust Inference on Functionals of Inverse Problems
- Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
- Reinforcement Learning in Low-Rank MDPs with Density Features
- LobsDICE: Offline Learning from Observation via Stationary Distribution Correction Estimation
- Playing Atari with Deep Reinforcement Learning
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Strong data processing inequalities and $\Phi$-Sobolev inequalities for discrete channels
- Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
- A Review of Off-Policy Evaluation in Reinforcement Learning
- A Researcher's Guide to Empirical Risk Minimization
- Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature Annealing
- Fitted Q-Evaluation without Bellman Completeness via Occupancy Weighting
- Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
- Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
- Inverse Reinforcement Learning with Just Classification and a Few Regressions
- What are the Statistical Limits of Offline RL with Linear Function Approximation?
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey