Convergence rates for the RMSprop optimizer with full control of the hyperparameters
cs.LG, math.OC, math.PR
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- Convergence of the Iterates for Momentum and RMSProp for Local Smooth Functions: Adaptation is the Key
- Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
- Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
- Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
- Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
- Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
- Convergence rates for the Adam optimizer
- Sharp higher order convergence rates for the Adam optimizer
- Global Stability and Step Size Robustness of RMSProp
- Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
- Non asymptotic analysis of Adaptive stochastic gradient algorithms and applications
- On the Convergence of Adam, Revisited
- Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks
- Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory
- Asymptotic Convergence and Stability of Adaptive Gradient Methods in Smooth Non-convex Optimization
- Adam: A Method for Stochastic Optimization
- Convergence of Adam Under Relaxed Assumptions
- Decoupled Weight Decay Regularization
- A Qualitative Study of the Dynamic Behavior for Adaptive Gradient Algorithms
- On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks