Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning
cs.LG, cs.SY, eess.SY
Submitted: 2026-09-20
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces.
Terminology
Abstract
Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.
Sources
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
- CAMMARL: Conformal Action Modeling in Multi Agent Reinforcement Learning
- CoBERL: Contrastive BERT for Reinforcement Learning
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks