Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
cs.LG, math.OC
Submitted: 2026-03-19
Updated: 2026-09-01
Comments: 35 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Convergence Analysis of a Momentum Algorithm with Adaptive Step Size for Non Convex Optimization
- An improvement of the convergence proof of the ADAM-Optimizer
- Non-Convergence and Limit Cycles in the Adam optimizer
- Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
- Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
- Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
- Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
- Convergence rates for the Adam optimizer
- ODE approximation for the Adam algorithm: General and overparametrized setting
- Sharp higher order convergence rates for the Adam optimizer
- Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
- Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
- A Simple Convergence Proof of Adam and Adagrad
- Blow up phenomena for gradient descent optimization methods in the training of artificial neural networks
- A Novel Convergence Analysis for Algorithms of the Adam Family
- Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
- Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory
- Adam: A Method for Stochastic Optimization
- Convergence of Adam Under Relaxed Assumptions
- On the Convergence of Adam and Beyond
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks