A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad and Muon
cs.LG
Submitted: 2026-04-19
Updated: 2026-08-27
Terminology
Sources
- SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine Variance
- The Geometry of Sign Gradient Descent
- Modular Duality in Deep Learning
- Old Optimizer, New Norm: An Anthology
- Beyond Uniform Smoothness: A Stopped Analysis of Adaptive SGD
- The duality structure gradient descent algorithm: analysis and applications to neural networks
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- An objective-function-free algorithm for nonconvex stochastic optimization with deterministic equality and inequality constraints
- A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization
- Cosmic shear with small scales: DES-Y3, KiDS-1000 and HSC-DR1
- Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
- A Note on the Convergence of Muon
- COSMOS: A Hybrid Adaptive Optimizer for Memory-Efficient Training of LLMs
- Training Deep Learning Models with Norm-Constrained LMOs
- SOAP: Improving and Stabilizing Shampoo using Adam
- Increased and Varied Radiation during the Sun's Encounters with Cold Clouds in the last 10 million years
- WNGrad: Learn the Learning Rate in Gradient Descent
- Structured Preconditioners in Adaptive Optimization: A Unified Analysis
- Block-Normalized Gradient Method: An Empirical Study for Training Deep Neural Network
- AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks