Second-Moment Stochastic Approximation Methods
math.OC, cs.AI, stat.ML
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Adam: A Method for Stochastic Optimization
- ASGO: Adaptive Structured Gradient Optimization
- Adam-mini: Use Fewer Learning Rates To Gain More
- The AdEMAMix Optimizer: Better, Faster, Older
- Quasi-hyperbolic momentum and Adam for deep learning
- MARS: Unleashing the Power of Variance Reduction for Training Large Models
- Adaptive Matrix Online Learning through Smoothing with Guarantees for Nonsmooth Nonconvex Optimization
- Stochastic Approximation with Block Coordinate Optimal Stepsizes
- Muon is Scalable for LLM Training
- The Llama 3 Herd of Models
- An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
- Why Gradients Rapidly Increase Near the End of Training
- Scale Weight Decay and Train Better
- The Newton-Muon Optimizer
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification