Optimizer choice matters for the emergence of Neural Collapse
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Optimizer choice matters for the emergence of Neural Collapse".
Jane: The paper was written by Jim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien Lucchi, University of Basel et al. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: The authors begin by grounding their entire argument in a massive scale experiment, running over three thousand nine hundred training runs across various datasets and architectures. This provides enormous empirical weight to the central claim that "Optimizer Choice Matters for the Emergence of Neural Collapse."
Jane: That sheer volume of data is what allows us to move beyond single case studies and really see how consistent this pattern is across different kinds of models, which is incredibly reassuring.
Lu: They are showing that while traditional views suggested NC was a universal constant, they found a clear distinction between the way weight decay is applied in adaptive optimizers.
Meng: Specifically, they're highlighting the difference between Adam and AdamW, which is such a practical distinction we rarely notice when deploying models in real-world scenarios.
Lalam: It’s fascinating to see that the subtle way weight decay is coupled or decoupled leads to fundamentally different training outcomes for AI models, moving from an assumption of inevitability to a nuanced reality.
Improvements/Metrics: Tom: The paper moves beyond just raw numbers and introducing a new diagnostic metric called NC0, which is defined as having zero row sum of last-layer weights. This gives us a much more precise tool for evaluation than the traditional metrics.
Jane: It’s a much more tractable way to measure the collapse than previous metrics, and it acts as a necessary condition for full Neural Collapse, meaning if NC0 diverges, we know right away that the model isn't collapsing fully.
Lu: The theoretical proofs they provide are quite impressive; demonstrating that under decoupled weight decay, NC0 does not vanish is a rigorous mathematical argument against the idea of universal collapse across optimization methods.
Meng: This gives us a much clearer signal to monitor during training, providing a clear indicator of whether the model is truly reaching that highly symmetric state or if it's just hitting an artificial plateau.
Lalam: The insights here suggest that understanding the mechanics of optimization isn't just an academic exercise; it’s a critical path for improving the reliability and predictability of AI systems by linking theoretical bounds to practical metrics.
Conclusion: Tom: So, after all this research, we have a clear picture: the way we apply weight decay—whether it’s coupled or decoupled—is crucial for whether Neural Collapse occurs. This really highlights the power of "Optimizer Choice Matters for the Emergence of Neural Collapse."
Jane: It's a massive correction to existing literature that NC is universal; we've learned that the optimizer is an active participant in shaping how AI learns.
Lu: The idea that momentum also accelerates NC, which they found first, adds another critical layer to this complex picture of optimization dynamics and how those layers interact.
Meng: This means our practical choice of tools for building these models has a direct impact on their convergence properties and will likely lead to more robust design choices going forward.
Lalam: I hope that the implications of "Optimizer Choice Matters for the Emergence of Neural Collapse" inspire a culture where we prioritize understanding how AI learns, not just that it achieves performance.
Conclusion: Tom: This entire discussion confirms that the fundamental way we set up our training processes—specifically how we handle weight decay—is not just a technical detail, but a defining factor in whether AI models experience Neural Collapse or not.
Jane: It's such an important distinction because it proves that NC isn't some unavoidable physical law of deep learning architecture; the way we’ve chosen to optimize is what fundamentally dictates the resulting model behavior.
Lu: The discovery that coupled weight decay enables this collapse while decoupled versions do not is a massive theoretical breakthrough, opening up new avenues for how we analyze the stability and geometry of large-scale AI models.
Meng: From a practical standpoint, this means that when designing new training pipelines for high-stakes AI systems, we can no longer treat optimizers as interchangeable; we need to account for their internal mechanics.
Lalam: This paper, "Optimizer choice matters for the emergence of Neural Collapse," teaches us that understanding the subtle interplay between how we train and what the resulting geometric structure looks like is a profound shift in how we view intelligence itself.
Tom: I think it really highlights that everything from SGD to AdamW has a different story to tell about their final convergence state.
Jane: And while they are showing that, it’s also encouraging to see the researchers explore the impact of momentum, as that adds another layer of complexity we need to understand.
Meng: It's a huge lesson for us in understanding the nuances of optimization—we can't just throw a large number at a weight decay parameter and expect a universal result.
Lu: We must also keep in mind that even if these findings are interesting, they don’re not the end of the story, and we have so many more questions about how these geometric properties manifest across different layers of larger models.
Tom: Exactly, it's a really deep rabbit hole to go down.
Jim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien Lucchi, University of Basel, Warsaw University of Technology, IDEAS Research Institute
cs.LG
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: Published as a conference paper at ICLR 2026
Code: https://github.com/rhubarbwu/neural-collapse
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper investigates the role of optimization algorithms in the emergence of Neural Collapse (NC), a phenomenon where deep neural network representations self-organize into highly symmetric
Key concepts
- Neural Collapse (NC)
- NC refers to a state where an AI model collapses into a highly symmetric state during training. The research findings suggest this phenomenon is not inevitable, but instead depends on the specific optimization setup used.
- Weight Decay
- This is the mechanism used during training to regulate how weights are updated. The study focuses on whether this decay is 'coupled' or 'decoupled,' noting that the subtle difference between these two methods significantly impacts the final training outcome.
- NC0
- NC0 is a specific diagnostic metric introduced in the paper. It measures whether the last-layer weights have a zero row sum, providing a precise and tractable way to evaluate if an AI model is truly experiencing full Neural Collapse.
Terminology
Summary
This paper investigates the role of optimization algorithms in the emergence of Neural Collapse (NC), a phenomenon where deep neural network representations self-organize into highly symmetric geometric structures during training. While previous theoretical works suggested NC might be universal across all optimization methods, this research challenges that assumption, demonstrating that the choice of optimizer—specifically how weight decay is implemented—is a critical determinant in whether NC occurs.
The Core Problem and New Metrics
Neural Collapse manifests through several geometric properties involving last-layer features and weights during the terminal phase of training (TPT). These include:
-
NC1 - Variability Collapse: within-class variability vanishes as features collapse to class means.
-
NC2 - Convergence of Centered Class Means to Simplex ETF: class means converge to a maximally symmetric configuration.
-
NC3 - Convergence to Self-Duality: alignment between last-layer weights and class means.
-
NC4 - Simplification to Nearest-Class-Center: decision boundaries simplify to those of a nearest-class-mean classifier.
The authors note that existing NC metrics are difficult to track and analyze theoretically
because they often plateau at small, non-zero values under realistic training regimes. To solve this, they introduce a novel diagnostic metric, NC0 (the zero row sum of the last-layer weight), asserting that its convergence to zero is a necessary condition for NC.
This allows for a more definitive assessment: if NC0 diverges, we can conclude that NC can not occur.
The Role of Weight Decay Coupling
The central finding is that the emergence of NC is highly dependent on whether weight decay is coupled
or decoupled.
Through extensive experiments involving over 3,900 training runs, the authors demonstrate that:
** Training with AdamW (which uses decoupled weight decay) does not lead to an NC solution. **
** Training with SGD or Adam (which use coupled weight decay) does lead to an NC solution. **
The paper identifies coupled weight decay as a key driver of NC in realistic settings.
In adaptive optimizers like Adam, the regularization is applied directly within the gradient update, whereas AdamW separates it. The authors show that this subtle yet consequential distinction has been largely overlooked in prior literature,
and they provide theoretical evidence that SignGD with decoupled weight decay fails to satisfy NC0, while SignGD with coupled weight decay can lead to it.
Theoretical and Empirical Insights
The researchers support their findings with several theoretical proofs regarding the convergence of the NC0 metric. They prove that for SGD, NC0 converges to zero at an exponential rate proportional to the weight decay.
Furthermore, they provide the first result concerning momentum in the context of NC, demonstrating an accelerating effect of momentum on NC
when training with SGD.
Empirical results across various datasets (MNIST, FashionMNIST, CIFAR10), architectures (ResNet9, VGG9), and optimizers confirm these dynamics. Notably, while AdamW and SignumW exhibit much larger NC metrics that remain strictly away from zero,
SGD and Adam achieve significantly lower values. The authors also observe partial neural collapse,
where some NC properties (like NC1 or NC2) might emerge even when others (like NC0 or NC3) fail to satisfy the full collapse criteria. Finally, they demonstrate that while coupled weight decay is essential for achieving strong geometric alignment, it is not strictly necessary for achieving strong generalization performance.
Improvements for AI systems
Based on the theoretical and empirical findings in this paper, I propose the following specific architectural and algorithmic improvements to deep learning training pipelines:
- Implementation of
Coupled Weight Decay
for Adaptive Optimizers
By transitioning from decoupled weight decay (AdamW) to coupled weight decay (Adam/Signum-style) during the terminal phase of training, we can force the model into a true Neural Collapse (NC) state.
-
The improved AI system will exhibit superior geometric stability in its last-layer representations.
-
It will achieve significantly higher self-duality between weights and class means, leading to more robust decision boundaries and better out-of-distribution (OOD) detection capabilities compared to standard AdamW models.
- Momentum-Accelerated Neural Collapse Tuning
Since the paper proves that momentum accelerates the convergence of NC metrics (specifically NC0) independently of training loss convergence, we can implement a dynamic momentum scheduling specifically optimized for geometric alignment.
-
The improved AI system will reach highly symmetric, optimal feature configurations in significantly fewer epochs than standard SGD or Adam.
-
It allows for a
fast-track
training regime where the model is identified as having reached its optimal geometric state based on NC0 convergence rather than waiting for loss plateaus, reducing wasted compute.
- NC0-Guided Training Diagnostics and Early Stopping
Replace traditional loss/accuracy monitoring with the proposed NC0 metric (Zero Row Sum of Last-Layer Weights) as a primary diagnostic tool.
-
The improved AI system will use NC0 to distinguish between
illusory collapse
(where metrics like NC3 plateau at non-zero values in AdamW) andtrue collapse
(where NC0 vanishes). -
This enables high-precision early stopping: training can be terminated the moment NC0 hits a predefined threshold, preventing overfitting and saving massive computational costs in large-scale training runs.
- Hybrid Optimizer Switching (AdamW to Coupled SGD/Adam)
Implement an automated scheduler that starts training with AdamW (for rapid initial convergence and high accuracy) and switches to a coupled weight decay regime (SGD or Adam with coupled WD) during the terminal phase of training.
-
The improved AI system will combine the fast convergence of decoupled adaptive methods with the superior geometric properties (Neural Collapse) of coupled methods.
-
This results in a model that achieves high generalization accuracy while maintaining the structural benefits of NC, such as improved transfer learning and better robustness to class imbalance.
Abstract
Neural Collapse (NC) refers to the emergence of highly symmetric geometric structures in the representations of deep neural networks during the terminal phase of training. Despite its prevalence, the theoretical understanding of NC remains limited. Existing analyses largely ignore the role of the optimizer, thereby suggesting that NC is universal across optimization methods. In this work, we challenge this assumption and demonstrate that the choice of optimizer plays a critical role in the emergence of NC. The phenomenon is typically quantified through NC metrics, which, however, are difficult to track and analyze theoretically. To overcome this limitation, we introduce a novel diagnostic metric, NC0, whose convergence to zero is a necessary condition for NC. Using NC0, we provide theoretical evidence that NC cannot emerge under decoupled weight decay in adaptive optimizers, as implemented in AdamW. Concretely, we prove that SGD, SignGD with coupled weight decay (a special case of Adam), and SignGD with decoupled weight decay (a special case of AdamW) exhibit qualitatively different NC0 dynamics. Also, we show the accelerating effect of momentum on NC (beyond convergence of train loss) when trained with SGD, being the first result concerning momentum in the context of NC. Finally, we conduct extensive empirical experiments consisting of 3,900 training runs across various datasets, architectures, optimizers, and hyperparameters, confirming our theoretical results. This work provides the first theoretical explanation for optimizer-dependent emergence of NC and highlights the overlooked role of weight-decay coupling in shaping the implicit biases of optimizers.
Sources
- On the Role of Neural Collapse in Transfer Learning
- Limitations of Neural Collapse for Understanding Generalization in Deep Learning
- An Unconstrained Layer-Peeled Perspective on Neural Collapse
- Generalized Neural Collapse for a Large Number of Classes
- Adam: A Method for Stochastic Optimization
- Detecting Out-of-Distribution Through the Lens of Neural Collapse
- SOAP: Improving and Stabilizing Shampoo using Adam
- Linguistic Collapse: Neural Collapse in (Large) Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks