Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

arXiv:2608.12973 · stat.ML, cs.LG · Submitted 2026-08-14 · Read on arXiv

Zijie Cheng, Yang Peng, Zhihua Zhang

Peking University · Tsinghua University

stat.ML, cs.LG

Submitted: 2026-08-14

Updated: 2026-08-17

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper studies statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning.

Terminology

Summary

This paper studies statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. The authors establish functional central limit theorems for both synchronous and asynchronous QTD, showing that the averaged iterates converge weakly to a rescaled Brownian motion. They then develop online inference procedures based on random scaling, constructing an asymptotically pivotal statistic that can be computed recursively without storing the entire trajectory of QTD iterates.

The paper assumes access to a generative model in a tabular γ-discounted Markov decision process. For a given policy π and quantile level m, the goal is to characterize the asymptotic distribution of QTD and develop valid online inference procedures based on the QTD iterates.

The main contributions are:

  1. Establishing functional central limit theorems in Theorem 3.1 and 3.2 for the iterates of both synchronous and asynchronous QTD. After suitable normalization, the partial-sum processes of the iterates converge weakly to rescaled Brownian motions, with explicit characterizations of the corresponding asymptotic covariance matrices.

  2. Developing inference procedures based on random scaling for QTD in Theorem 3.3. The authors construct an asymptotically pivotal statistic, avoiding the need to explicitly estimate the asymptotic covariance matrices. The quantities in the statistic can be updated recursively along the QTD trajectory, so the procedures can be implemented online without storing the entire sequence of iterates.

The paper notes that answering these questions for QTD is challenging due to the unique structure of quantile projected Bellman operators. Unlike categorical distributional reinforcement learning where the categorical-projected Bellman update acts linearly on probability vectors, QTD leads to a nonlinear and non-smooth stochastic approximation recursion, where the stochastic updates are generated by indicator-type transformations of the quantile parameters. This non-smooth dependence across different quantile levels prevents a direct application of existing stochastic approximation inference theory.

The theoretical results rely on two main assumptions:

  • Assumption 1: For any state-action pair, the reward distribution is supported on [0,1] and has a Lipschitz continuous Lebesgue density that is bounded above and strictly positive on (0,1).

  • Assumption 2: For every pair of states with positive transition probability, the quantile locations satisfy θm(s,i) - γθm(s',j) ∉ 0,1 for all quantile indices, excluding boundary-degenerate configurations.

The functional central limit theorem for synchronous QTD (Theorem 3.1) states that as T → ∞, the normalized partial-sum process converges weakly to G−1Γ syn(1/2)B(·), where G is the Jacobian matrix of the mean-field function at the fixed point, Γ syn is the covariance matrix of the noise, and B(·) is a standard Brownian motion. For asynchronous QTD (Theorem 3.2), the limit is (D μG)−1Γ asyn(1/2)B(·), where D μ is a diagonal matrix of the sampling distribution over states.

The online inference procedure (Theorem 3.3) constructs a random scaling matrix V̂ T that can be updated recursively. The resulting statistic (c J(θ̄ T - θ m))/√(c J V̂ T c) converges in distribution to a pivotal random variable V = B(1)/∫01(B(u) - uB(1))2du, which is independent of the asymptotic covariance matrices. This leads to an asymptotic confidence interval in Corollary 3.1.

The proof strategy involves applying a general functional central limit theorem (Theorem 4.1) for stochastic approximation recursions with non-smooth mean-field functions. The authors verify the required conditions by showing that the mean-field function h has a unique zero at θ m, that its Jacobian G is uniformly repulsive, and that the noise sequence satisfies the martingale difference, boundedness, and conditional covariance convergence conditions.

The paper concludes by suggesting future directions, including generalizing results to more general step size schedules and establishing non-asymptotic convergence rates for QTD to better understand finite-sample behavior.

Improvements for AI systems

Improvements to AI Systems:

  1. Online Uncertainty Quantification for Distributional RL Agents
  • Improvement: Integrate the recursive random-scaling statistic (Theorem 3.3) into QTD-based agents (e.g., IQN, FQF) to produce real-time confidence intervals for estimated quantiles without storing full iterate histories.

  • Capability: The AI can now report statistically valid uncertainty bounds during training or deployment, enabling risk-aware decision-making (e.g., safe exploration, anomaly detection) in live environments with minimal memory overhead.

  1. Non-Smooth Stochastic Approximation Inference
  • Improvement: Generalize the functional central limit theorem (Theorem 4.1) to other RL algorithms with non-smooth, indicator-based updates (e.g., quantile regression in actor-critic methods, ranking losses).

  • Capability: AI systems can now perform hypothesis testing and construct confidence intervals for parameters learned via non-differentiable objectives, extending beyond QTD to broader non-smooth optimization in RL.

  1. Asynchronous Sampling-Aware Variance Estimation
  • Improvement: Use the asynchronous QTD covariance characterization (Theorem 3.2) to correct for state-visitation bias in off-policy or replay-buffer settings.

  • Capability: The AI can automatically adjust its inference to account for non-uniform sampling distributions, producing unbiased uncertainty estimates even when data collection is imbalanced—critical for real-world robotics or recommendation systems.

  1. Automatic Degeneracy Detection
  • Improvement: Leverage Assumption 2 to design a pre-training check that flags boundary-degenerate quantile configurations (where θ m(s,i) - γθ m(s',j) ∈ 0,1).

  • Capability: The AI can proactively detect and avoid unstable learning regimes, reparameterizing or adjusting the quantile grid to ensure valid statistical inference and prevent divergence.

  1. Efficient Sequential Testing in RL
  • Improvement: Apply the pivotal statistic (V) to create an online sequential test for policy convergence or distributional shift, updating the test statistic recursively at each step.

  • Capability: The AI can halt training early when the quantile estimates are statistically stable (e.g., V falls within a pre-specified threshold), saving compute while guaranteeing a user-defined confidence level.

  1. Transferable Confidence Intervals for Offline RL
  • Improvement: Adapt the covariance matrix formulas (Γ syn, Γ asyn) to estimate uncertainty in offline QTD from a fixed dataset, using the generative-model assumption as a proxy for empirical transition probabilities.

  • Capability: The AI can provide calibrated confidence intervals for value estimates in offline RL, enabling safer deployment in medical or financial settings where online interaction is costly or risky.

  1. Adaptive Step-Size Scheduling with Inference Guarantees
  • Improvement: Extend the FCLT results to time-varying step sizes (as suggested in future work) and implement an adaptive scheduler that adjusts step size based on the online pivotal statistic.

  • Capability: The AI can dynamically trade off convergence speed and inference accuracy, automatically slowing down when uncertainty is high to maintain valid confidence intervals—useful for lifelong learning agents.

Sources

Related papers