Double Descent and Malign Overfitting in Diffusion Models
cs.LG, cond-mat.dis-nn
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 44 pages, 17 figures
Code: https://github.com/mseitzer/pytorch-fid
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conventional wisdom in deep learning holds that overparameterization---having more parameters p than training samples n ---is benign: larger models generalize better and, even without regularization,
Terminology
Abstract
Conventional wisdom in deep learning holds that overparameterization---having more parameters p than training samples n ---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number m of noise realizations per training sample, an interpolation peak does occur, but at p about nm rather than at p about n as in standard regression. The rise of the test loss, however, sets in much earlier, at p about n, independently of m. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at p about n; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with m 1, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.
Sources
- Losing dimensions: Geometric memorization in generative diffusion
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Wider Networks Learn Better Features
- Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime
- A Good Score Does not Lead to A Good Generative Model
- Understanding diffusion models requires rethinking (again) generalization
- Generalization Dynamics of Linear Diffusion Models
- Optimal Regularization Can Mitigate Double Descent
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Manifolds, Random Matrices and Spectral Gaps: The geometric phases of generative diffusion
- Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks