Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

arXiv:2609.23775 · cs.LG, cs.SY, eess.SY · Submitted 2026-09-20 · Read on arXiv

cs.LG, cs.SY, eess.SY

Submitted: 2026-09-20

Updated: 2026-09-20

License: http://creativecommons.org/licenses/by/4.0/

The gist: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces.

Terminology

Abstract

Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.

Sources

Related papers