Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
stat.ML, cs.LG
Submitted: 2026-05-05
Updated: 2026-09-11
Comments: 45 pages, 11 figures, 1 table
License: http://creativecommons.org/licenses/by/4.0/
The gist: We provide a theoretical analysis of Adam under non-stationary stochastic objectives, separating two regimes: Euclidean tracking under adaptive strong monotonicity of the Adam-preconditioned
Terminology
Abstract
We provide a theoretical analysis of Adam under non-stationary stochastic objectives, separating two regimes: Euclidean tracking under adaptive strong monotonicity of the Adam-preconditioned mean-gradient operator, and high-probability projected stationarity guarantees under general L-smooth objectives. In the tracking regime, we derive finite-time expected and high-probability bounds that decompose sharply into four components: initialization, objective drift, a first-moment tracking error governed by β 1, and a preconditioner perturbation governed by β 2. We characterize the burn-in time required for the transient terms to decay to the asymptotic tracking bound under constant and step-decay schedules. We also prove a high-probability bound on the average projected stationarity gap for Adam under distribution shift. Across both analyses, our bounds reveal a noise--drift tradeoff: in noise-dominated regimes, first-moment averaging and adaptive preconditioning can yield favorable upper guarantees, whereas in drift-dominated regimes, stale first-moment information and preconditioner perturbations can enlarge Adam's tracking guarantee, potentially allowing vanilla SGD to attain a smaller tracking error. Our explicit (β 1,β 2,ε) -dependent bounds identify mechanisms through which adaptive step-sizing can help or hurt under nonstationarity and provide theoretical explanations consistent with Adam's empirical instability and stabilization under distribution shift.
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Maintaining Plasticity in Deep Continual Learning
- Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
- On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
- SGD with Dependent Data: Optimal Estimation, Regret, and Inference
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey