Beyond Quadratic Loss: The Stability Phase Diagram of Adam
cs.LG, math.OC, stat.ML
Submitted: 2026-09-16
Updated: 2026-09-25
Comments: 20 pages, 9 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms.
Terminology
Abstract
Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the (β 1,β 2) plane. Across a range of model--task settings, an approximately linear boundary, 1-β 2=C(1-β 1), separates spiky from non-spiky dynamics, whereas a one-dimensional quadratic loss produces approximately cubic slope. A one-dimensional superquadratic loss L(x) proportional tox n recovers the near-linear scaling and links the boundary coefficient to the effective loss exponent n. We further show that confident cross-entropy losses develop a core--wall landscape comprising a narrow quadratic core followed by a steep wall, which produces effective superquadratic behavior at the scale of an optimizer update. Together, these results connect Adam loss spikes to both the mismatch between momentum timescales and finite-scale superquadratic loss geometry beyond the Hessian.
Sources
- Towards Understanding Adam Convergence on Highly Degenerate Polynomials
- Adaptive Preconditioners Trigger Loss Spikes in Adam
- Unveiling the Basin-Like Loss Landscape in Large Language Models
- Adaptive Gradient Methods at the Edge of Stability
- Understanding Optimization in Deep Learning with Central Flows
- Adam: A Method for Stochastic Optimization
- Visualizing the Loss Landscape of Neural Nets
- Loss Spike in Training Neural Networks
- Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay
- Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
- A Theory on Adam Instability in Large-Scale Machine Learning
- Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- The Slingshot Mechanism: An Empirical Study of Adaptive Optimizers and the Grokking Phenomenon
- Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks