Fast Best-in-Class Regret for Contextual Bandits
Samuel Girard, Aurelien Bibaut, Arthur Gretton, Nathan Kallus, Houssam Zenati
stat.ML, cs.LG
Submitted: 2026-08-19
Updated: 2026-08-20
Terminology
Sources
- Efficient Optimal Learning for Contextual Bandits
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Sequential Off-Policy Learning with Logarithmic Smoothing
- Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality
- Second Order Bounds for Contextual Bandits with Function Approximation
- Estimating means of bounded random variables by betting
- Policy Learning with Adaptively Collected Data
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey