Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling
stat.ML, cs.IT, cs.LG, math.IT, math.ST, stat.TH
Submitted: 2026-09-02
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study a variant of the Thompson Sampling (TS) algorithm, called α-TS, for solving stochastic generalized linear bandit problems.
Terminology
Abstract
We study a variant of the Thompson Sampling (TS) algorithm, called α-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing α-TS that uses a fractional or α-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of α-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of α proportional to d-1, our general regret bound yields the best known regret bound of O(d 3/2 sqrt T T) for both the exponential and sub-Gaussian families of reward distributions. We further provide an α-dependent lower bound showing that the regret constant depends on the product αd, and that when α proportional to d-1 the regret scales as Ω(d 3/2 sqrt T), explaining the origin of the d 3/2 factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.
Sources
- Thompson Sampling for Contextual Bandits with Linear Payoffs
- The Elliptical Potential Lemma Revisited
- On Frequentist Regret of Linear Thompson Sampling
- Geometry-Aware Approaches for Balancing Performance and Theoretical Guarantees in Linear Bandits
- Feel-Good Thompson Sampling for Contextual Bandits and Reinforcement Learning
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey