Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions
stat.ML, cs.LG, math.OC, stat.ME
Submitted: 2024-06-12
Updated: 2026-09-07
Comments: Extended version; supersedes conference paepr
License: http://creativecommons.org/licenses/by/4.0/
The gist: Offline reinforcement learning enables evaluation and optimization of sequential decisions from historical data, when it is not possible to deploy new policies online due to safety, cost, and other
Terminology
Abstract
Offline reinforcement learning enables evaluation and optimization of sequential decisions from historical data, when it is not possible to deploy new policies online due to safety, cost, and other concerns. Big data advances enable rich state information, but may naively include reward- and action- irrelevant dynamics that are ultimately unnecessary for learning optimal actions. We introduce state abstractions that target preservation of the difference-of-Q functions, and we propose to learn these abstractions via causal machine learning of the difference-of-Q function and standard statistical sparse learning. Under a nonparametric additive-rewards model, we characterize when decision-centered abstractions are simpler than the full state space, motivating our estimation procedure. We develop a dynamic generalization of the R learner (Nie et al. 2021, Lewis and Syrgkanis 2021) for estimating difference of Q-functions, for discrete-valued actions a, a0. We leverage orthogonal estimation to improve convergence rates, even if the required estimates of Q and behavior policy converge at slower rates and prove consistency of policy optimization under a margin condition. The method can leverage black-box estimators of the Q-function and behavior policy to target estimation of a more structured Q-function contrast, and uses simple squared-loss minimization. We demonstrate variance improvements from our estimator and how our approach enables us to isolate the information needed for sequential decision-making, which can be less than that for state prediction, in simulated data and simulator-augmented real data.
Sources
- OpenAI Gym
- Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
- Orthogonal Statistical Learning
- Off-policy Evaluation with Deeply-abstracted States
- Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes
- Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning
- Towards optimal doubly robust estimation of heterogeneous causal effects
- Semiparametric doubly robust targeted double machine learning: a review
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Double/Debiased Machine Learning for Dynamic Treatment Effects via g-Estimation
- On Weighted Orthogonal Learners for Heterogeneous Treatment Effects
- Skill or Luck? Return Decomposition via Advantage Functions
- Statistically Efficient Advantage Learning for Offline Reinforcement Learning in Infinite Horizons
- Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning
- Denoised MDPs: Learning World Models Better Than the World Itself
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey