Superposition Yields Robust Neural Scaling
Yizhou Liu, Ziming Liu, Jeff Gore
Massachusetts Institute of Technology
cs.LG, cs.AI, cs.CL
Submitted: 2025-11-29
Updated: 2026-08-11
Comments: Best Paper Runner-up at NeurIPS 2025
Journal ref: Advances in Neural Information Processing Systems 38 (2025) 159269--159305
DOI: 10.52202/085713-5320
Code: https://github.com/liuyz0/SuperpositionScaling
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
The gist: The paper "Superposition Yields Robust Neural Scaling" proposes that "representation superposition, meaning that LLMs represent more features than they have dimensions, can be a key contributor to
Terminology
Summary
The paper Superposition Yields Robust Neural Scaling
proposes that representation superposition, meaning that LLMs represent more features than they have dimensions, can be a key contributor to loss and cause neural scaling.
To investigate this, the authors utilize a toy model based on Anthropic’s autoencoder, where they use weight decay to control the degree of superposition, allowing us to systematically study how loss scales with model size.
The research distinguishes between two primary regimes of superposition:
1. Weak Superposition Regime
In this regime, the model only represents the most frequent features and ignores others. The authors find that the scaling of loss with model dimension depends sensitively on how feature frequency decays with rank: the loss follows a power law with model size only if the feature frequencies themselves follow a power law.
Specifically, if feature frequencies follow a power law (p i proportional to 1/i alpha with alpha > 1), the loss or the summation starting at m will be a power law with m with exponent alpha - 1.
In this regime, loss in the weak superposition regime depends on summing frequencies of ignored features.
2. Strong Superposition Regime
In contrast, under strong superposition, the loss generically scales inversely with model dimension across a broad class of frequency distributions, due to geometric overlaps between representation vectors.
The authors explain that loss arises from the non-zero overlaps between representation vectors.
For even feature frequencies, the squared overlaps scaling like 1/m when n m, leading to the robust 1/m power-law loss.
They also note that for skewed feature frequencies (large alpha), the model exponent alpha m increases, with a conjectured worst-case scaling of alpha m about 2(alpha - 1).
Application to Large Language Models (LLMs)
The authors confirm that open-sourced LLMs operate in the strong superposition regime and have loss scaling inversely with model dimension.
Their analysis of the language model head weight matrix W reveals that the mean squared overlaps of normalized rows W i / W i 2 roughly obey 1/m scaling.
Furthermore, they find that token frequencies... follow a power law with exponent alpha close to 1.
The empirical model exponent for the studied LLMs was found to be alpha m = 0.91 plus or minus 0.04, which is close to 1.
The authors state that the Chinchilla scaling laws are also consistent with this behavior,
noting that given the relationship between model size N and dimension m (N proportional to m 2.52 plus or minus 0.03), the predicted exponent alpha m is 0.88 plus or minus 0.06.
In conclusion, the results identify representation superposition as a central driver of neural scaling laws,
providing a mechanistic explanation for the observed power-law decay of loss in large-scale machine learning systems.
Improvements for AI systems
1. Overlap-Minimizing Regularization (OMR)
- What it does: Replaces standard L 2 weight decay with a specialized loss term that penalizes the mean squared overlap of normalized representation vectors. This forces the model to maintain the 1/m geometric scaling of overlaps, minimizing the interference-driven loss that occurs when features collide in the same subspace.
2. alpha-Adaptive Architecture Scaling
- What it does: Automatically determines the optimal model dimension (m) and parameter count (N) for a specific dataset by first calculating the power-law exponent (alpha) of the dataset's token frequencies. This ensures the model is architecturally tuned to operate in the
strong superposition
regime, preventing the inefficient loss scaling characteristic of theweak superposition
regime.
3. Superposition-Optimized Weight Initialization
- What it does: Initializes the language model head and hidden layers such that the initial geometric overlap of representation vectors follows the 1/m scaling law. This allows the model to bypass the unstable early training phases and immediately leverage robust neural scaling laws.
4. Interference-Aware Feature Pruning
- What it does: Instead of pruning weights based on magnitude, this method prunes dimensions based on their contribution to
geometric overlap noise.
This creates highly compressed, efficient models that retain a high feature-to-dimension ratio without the performance degradation caused by feature interference.
5. Dynamic Feature-Density Tuning (DFDT)
- What it does: During training, this system monitors the relationship between model dimension and loss. If the model drifts toward the
weak superposition
regime (where loss depends heavily on feature frequency decay), the system dynamically adjusts the learning rate or weight decay to push the representations back into thestrong superposition
regime for more robust scaling.
Abstract
The success of today's large language models (LLMs) depends on the observation that larger models perform better. However, the origin of this neural scaling law, that loss decreases as a power law with model size, remains unclear. We propose that representation superposition, meaning that LLMs represent more features than they have dimensions, can be a key contributor to loss and cause neural scaling. Based on Anthropic's toy model, we use weight decay to control the degree of superposition, allowing us to systematically study how loss scales with model size. When superposition is weak, the loss follows a power law only if data feature frequencies are power-law distributed. In contrast, under strong superposition, the loss generically scales inversely with model dimension across a broad class of frequency distributions, due to geometric overlaps between representation vectors. We confirmed that open-sourced LLMs operate in the strong superposition regime and have loss scaling inversely with model dimension, and that the Chinchilla scaling laws are also consistent with this behavior. Our results identify representation superposition as a central driver of neural scaling laws, providing insights into questions like when neural scaling laws can be improved and when they will break down.
Sources
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Scaling Laws for Autoregressive Generative Modeling
- GPT-4 Technical Report
- How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges
- Galactica: A Large Language Model for Science
- Evaluating Large Language Models Trained on Code
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- A Solvable Model of Neural Scaling Laws
- Learning Curve Theory
- Physics of Skill Learning
- Scaling Laws and Interpretability of Learning from Repeated Data
- Neural Scaling Laws Rooted in the Data Distribution
- A Resource Model For Neural Scaling Law
- Decoupled Weight Decay Regularization
- Neural Collapse: A Review on Modelling Principles and Generalization
- OPT: Open Pre-trained Transformer Language Models
- Qwen2.5 Technical Report
- Pointer Sentinel Mixture Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks