Double descent is the principle of least action
cs.LG, cs.AI, math.ST, physics.comp-ph, physics.data-an, stat.TH
Submitted: 2026-09-16
Updated: 2026-09-17
Comments: 11 pages, 2 figures, 1 table
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The test error of a model plotted against its number of parameters d falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon.
Terminology
Abstract
The test error of a model plotted against its number of parameters d falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature T, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the d degrees of freedom in shares of T/2, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the L squared norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing d, effectively increasing weight regularization.
Sources
- Layer Normalization
- An Introduction to Flow Matching and Diffusion Models
- Adam: A Method for Stochastic Optimization
- Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning
- Decoupled Weight Decay Regularization
- Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
- Deep Double Descent: Where Bigger Models and More Data Hurt
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks