A Defense of the Quadratic Model
Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian
cs.LG, cs.AI, math.OC, stat.ML
Submitted: 2026-07-23
Code: https://github.com/alexandrumeterez/QuadraticModel
License: http://creativecommons.org/licenses/by/4.0/
The gist: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is,
Terminology
Abstract
Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model of optimization -- the quadratic model -- and show that it can be surprisingly predictive in an LLM setting with 150M parameters and 3B training tokens. Specifically, we show that Taylor expanding the model and the loss function at intermediate checkpoints through training can accurately predict the optimization dynamics over windows that can last up to 10% of training. Having established this agreement, we then turn to analyzing the structure of these local quadratic optimization problems through two lenses: the Hessian spectrum and local stability. Using Lanczos quadrature with extremely deep probes, we are able to estimate the Hessian spectrum deep into the tail, and we find a surprising amount of structure in both the eigenvalues and eigenvectors, which depends on the batch size, preconditioner, and training time. We also empirically test local linear stability at intermediate checkpoints and compare it to theoretical predictions to demonstrate that optimization in LLMs typically occurs at a stochastic edge of stability, whose nature is also determined by batch size. Our results indicate the quadratic model may be a theoretically tractable proxy for pretraining optimization dynamics.
Sources
- Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
- The proximal point method revisited
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Estimating the Spectral Density of Large Implicit Matrices
- On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Deep Curvature Suite
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- The large learning rate phase of deep learning: the catapult mechanism
- Neural Networks as Kernel Learners: The Silent Alignment Effect
- Learning Curves for SGD on Structured Features
- Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
- Properties of the After Kernel
- Randomized matrix-free quadrature: unified and uniform bounds for stochastic Lanczos quadrature and the kernel polynomial method
- Adaptive Gradient Methods at the Edge of Stability
- Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability
- A view of mini-batch SGD via generating functions: conditions of convergence, phase transitions, benefit from negative momenta
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks