On the Convergence of Adam, Revisited
Steven Heilman, Sampad Mohanty
cs.LG, math.OC, stat.ML
Submitted: 2026-07-03
Comments: 26 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-V3 Technical Report
- Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
- Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
- Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
- Muon Does Not Converge on Convex Lipschitz Functions
- Divergence of the ADAM algorithm with fixed-stepsize: a (very) simple example
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Divergence Results and Convergence of a Variance Reduced Version of ADAM
- Adam Converges Without Any Modification On Update Rules
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks