The Curse of Depth in Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Curse of Depth in Large Language Models".
Jane: The paper was written by Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng et al. from Westlake University, China and Emory University, USA and Dalian University of Technology, China and University of Surrey, UK and University of Oxford, UK.
Tom: Stay tuned as we take you through the paper and discuss its implications.
The Core Problem: Tom: So, "The Curse of Depth in Large Language Models" confirms that this problem isn's just an isolated bug; the phenomenon of ineffective deep layers exists across major LLM families like Llama and DeepSeek. This is a widespread issue, not a niche problem.
Jane: And the paper provides some really clear evidence, showing things like performance drop tests where removing one layer after its depth shows very little change in model performance at all levels. It’s shocking how resilient those deeper layers are.
Meng: If pruning deep layers barely hurts the overall accuracy, it strongly suggests that these layers aren't contributing much to the final output, which is a massive practical hurdle for optimization algorithms.
Lu: The paper's analysis points to a specific mechanism: Pre-Layer Normalization (Pre-LN). This approach is designed to stabilize training but doesn'output variance explodes exponentially with depth.
Lalam: That exponential growth of variance, as the model gets deeper, means the layers are essentially collapsing into an identity function. They aren't performing meaningful transformations anymore; they are just passing information along without changing it much.
Tom: The paper proves that when this happens, the derivative of those deep blocks approaches identity matrix. It’s mathematically proven that's why they become so ineffective.
Meng: This is where the efficiency problem really hits the fan, Tom. If we are paying for layers to behave like a simple pass-through function, we have to rethink how these massive architectures are constructed entirely differently.
The Solution: Tom: Thankfully, "The Curse of Depth in Large Language Models" suggests a clean fix: LayerNorm Scaling, or LNS. This is the proposed solution to mitigate the exponential variance growth.
Jane: LNS scales the output of layer normalization by /sqrt based on its depth. It’s a simple modification, but it’s incredibly powerful because it actively controls that variance explosion before it can happen.
Meng: I appreciate that LNS is hyperparameter-free and doesn't introduce extra parameters during training. That makes implementation extremely straightforward for large-scale deployment.
Lu: The theoretical implications are huge, Lu thinks, because by slowing the growth of variance from exponential to at most quadratic, we’re fundamentally changing the landscape of what a deep neural network can achieve. It expands our theoretical capacity dramatically.
Lalam: LNS helps ensure that every single layer contributes something unique and different to the final representation. Instead of just stacking redundant functions, we get a truly diverse set of features from every layer in the model.
Tom: Exactly, Lalam; it makes sure those deep layers are contributing meaningfully again. This approach is designed to counteract that identity mapping behavior and revitalize the contribution of those later blocks.
The Results: Tom: We’ve seen how the problem manifests, but now we’re looking at the empirical evidence in "The Curse of Depth in Large Language Models." The experiments show LNS consistently outperforms all other existing normalization techniques.
Jane: It shows a substantial improvement in both pre-training performance and when we move to supervised fine-tuning tasks. The model is learning better representations across different sizes, from 130M up to 7B parameters.
Meng: And the results are scalable; LNS works whether you're training a small Llama variant or a massive state-of-the-art OLMo model. It consistently delivers lower perplexity and more stable training dynamics for large systems.
Lu: The fact that the angular distance between layers increases under LNS also confirms that the representation is truly diversifying, not just because the variance is controlled, but because of how the signals are propagating through depth.
Lalam: This means we can finally build AI systems where every single step in every layer matters. It's a huge win for resource efficiency and quality simultaneously.
Tom: It’s clear that by fixing the mechanism of Pre-LN, we’ve successfully tackled the Curse of Depth and found a robust alternative.
Wrap Up: Tom: As we wrap up our discussion on "The Curse of Depth in Large Language Models," it seems like we've seen a major breakthrough in how these massive models are built and trained.
Jane: It’s genuinely exciting that a simple scaling factor, based on the square root of depth, can unlock so much more performance potential across different models.
Lu: I think the biggest intellectual takeaway is that we have successfully moved away from an exponential constraint toward a manageable polynomial growth in variance. That's a huge shift in our understanding of deep learning theory.
Meng: From an engineering viewpoint, this means that we can move forward with these architectures knowing that the performance gains aren't just theoretical; they are practical and scalable across massive training budgets.
Lalam: I hope this work encourages more leads to build truly efficient AI systems for the future, ensuring that every computational step serves a meaningful purpose.
Tom: Absolutely, Lalam. It’s clear that "The Curse of Depth in Large Language Models" offers a robust and computationally efficient path forward for the next generation of AI. Goodbye everyone, and we'll see you on the next paper!
Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu
Westlake University, China · Emory University, USA · Dalian University of Technology, China · University of Surrey, UK · University of Oxford, UK
cs.LG, cs.AI
Submitted: 2026-08-22
Updated: 2026-08-25
Importance score: 95/100
The gist: The paper analyzes "The Curse of Depth in Large Language Models," examining variance propagation and performance degradation as model depth increases.
Key concepts
- Curse of Depth
- A widespread issue in large language models where deep layers become ineffective. This occurs because the exponential growth of variance causes the layers to behave like an identity function, passing information without meaningful transformation.
- Pre-Layer Normalization (Pre-LN)
- A training approach designed to stabilize model training. However, it is susceptible to variance exploding exponentially with depth, which leads to the deep layers becoming redundant and less effective for computation.
- LayerNorm Scaling (LNS)
- The proposed solution that mitigates exponential variance growth by scaling the layer normalization output by 1/sqrt(ℓ), where ℓ is the depth. This simple modification stabilizes training and enhances performance.
Terminology
Summary
The paper analyzes The Curse of Depth in Large Language Models,
examining variance propagation and performance degradation as model depth increases.
Variance Growth and Training Dynamics:
Regarding the stability of training, the analysis tracks Variance Growth in Pre-LN Training.
Figure 9 illustrates this layer-wise variance for LLaMA-130M at 1000, 3000, and 6000 epochs. The findings show that In all cases, the variance follows an exponential growth pattern as depth increases, indicating that deeper layers experience uncontrolled variance amplification regardless of training progress.
This confirms that this issue persists throughout training rather than being a temporary effect,
necessitating stabilization techniques like LayerNorm Scaling.
Curse of Depth in Multimodal Models:
To test if this phenomenon affects vision-language models (VLMs), layer–pruning experiments were conducted on Qwen 2.5-VL-7B. Figure 10 presents the performance drop results, revealing a clear distinction between model components: We observe that the language branch clearly suffers from the Curse of Depth, whereas the vision branch remains uniformly important.
This suggests that the phenomenon is more pronounced in autoregressive language components of VLMs and may not directly transfer to vision encoders.
Theoretical Analysis and Gradient Bounding:
The paper provides a detailed theoretical analysis concerning gradient propagation. The attention mechanism is defined as Attn(x) = head 1,, head h W O. The core mathematical effort involves bounding the partial derivative d y L / d W 1.
The derivation establishes bounds for the gradient propagation. Specifically, an upper bound for d y L / d W 1 is found:
d y L over d W 1 1 over s+2+ sqrt s sqrt p over p cubed d + d head + sigma x l d head over h
By combining these bounds, the final upper bound for d y L / d W 1 is presented in the form:
d y L over d W 1 1 + A + 2 B
where A and B are complex terms involving various model parameters (sigma x'l, sigma x l, d, d head, etc.) and the layer index.
Limitations:
The authors acknowledge several limitations of their work:
-
Scope of Architectures:
Our study primarily focuses on Transformer-based LLMs using Pre-LN.
The theoretical study "does not cover models employing alternative normalization strategies (e.g., Post-LN only (Du et al., 2021), normalization-free architectures (Zhu et al., 2025a)) or emerging paradigms such as mixture-of-experts or structured sparsity-based models." -
Task Coverage:
Most empirical evaluations, including pruning and angular distance analyses, were conducted using general-purpose benchmarks like MMLU.
They note thatdomain-specific or long-context reasoning tasks may reveal different dynamics in deep layer contributions, which we leave for future work.
-
Fine-grained Representation Quality: While LayerNorm Scaling (LNS) improves angular distance and performance sensitivity, a deeper analysis of the information preserved is lacking. For example,
whether LNS helps preserve syntactic, semantic, or factual knowledge across depth is unclear.
Improvements for AI systems
Based on a rigorous analysis of this scientific paper, the following specific improvements can be implemented in state-of-the-art AI systems:
-
Mandatory Variance Control Mechanism: Implement LayerNorm Scaling (LNS) as the core normalization strategy within all transformer blocks. This involves applying a scaling factor inversely proportional to the square root of the layer index (sqrt 1/) to every Layer Normalization output, preventing uncontrolled exponential variance growth during training.
-
Targeted Pre-Training Optimization: Systematically replace standard Pre-Layer Normalization (Pre-LN) with LNS across all model sizes (e from 130M up to 7B). This ensures that the entire training regimen is optimized not just for stability, but for effective contribution from deep layers.
-
Initialization Alignment: When adopting LNS, strictly enforce the recommendation to remove Scaled Initialization. This prevents conflicting scaling mechanisms from diminishing the effectiveness of LNS, ensuring a clean and maximally effective implementation of the variance control strategy.
4 Specific Application in Vision Transformers (ViT): For multimodal architectures (e.g., ViT), apply LNS specifically after the Attention and MLP blocks to stabilize the forward signal, as empirical evidence shows optimal placement varies by architecture.
The resulting AI system, leveraging LayerNorm Scaling, will exhibit the following quantifiable improvements:
A. Enhanced Training Efficacy and Stability:
-
Mitigation of Identity Mapping Collapse: Deep layers will no longer behave as near-identity mappings (d y L / d x 1 to M). Instead, they will maintain measurable gradient contributions, allowing the model to utilize its full depth capacity.
-
Guaranteed Convergence and Stability: The system achieves significantly reduced loss spikes during training due to controlled variance growth. This allows for stable convergence even in large-scale pre-training (e.g., 20B tokens on OLMo).
-
Robustness across Model Scales: The improved system maintains optimal performance stability regardless of scaling up to 7 Billion parameters, avoiding the divergence issues observed in other advanced normalization techniques (like Mix-LN).
B. Superior Representation Quality and Generalization:
-
Increased Feature Diversity: Deep layers will generate significantly more distinct representations (higher angular distance from adjacent layers) compared to standard Pre-LN models, indicating a more rich and complex feature space.
-
Optimized Downstream Transfer: The enhanced diversity in deep layer representations translates directly into superior performance during Supervised Fine-Tuning (SFT), leading to higher accuracy across diverse benchmarks (e.g., MMLU, ARC-e).
C. Operational Efficiency:
-
Hyperparameter-Free Design: Unlike competing methods, LNS requires zero hyperparameter tuning, simplifying deployment and ensuring consistent performance across a highly variable set of architectures.
-
Minimal Overhead: The implementation introduces no additional learnable parameters or complex computational overhead beyond the original design of existing LayerNorm layers.
Sources
- GPT-4 Technical Report
- Layer Normalization
- Qwen2.5-VL Technical Report
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- GLM: General Language Model Pretraining with Autoregressive Blank Infilling
- Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
- OLMo: Accelerating the Science of Language Models
- The Unreasonable Ineffectiveness of the Deeper Layers
- GradientStabilizer:Fix the Norm, Not the Gradient
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
- Mistral 7B
- The Remarkable Robustness of LLMs: Stages of Inference?
- Outlier-weighed Layerwise Sampling for LLM Fine-tuning
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
- OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- GLU Variants Improve Transformer
- Spike No More: Stabilizing the Pre-training of Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks