The Curse of Depth in Large Language Models
summary
The gist
The paper analyzes "The Curse of Depth in Large Language Models," examining variance propagation and performance degradation as model depth increases.
In short
The episode discusses 'The Curse of Depth in Large Language Models,' a problem where deep layers become ineffective due to exponential variance growth. The hosts introduce LayerNorm Scaling (LNS), a simple fix that controls this variance, significantly improving model performance and efficiency across various LLM sizes.
Key concepts
- Curse of Depth
- A widespread issue in large language models where deep layers become ineffective. This occurs because the exponential growth of variance causes the layers to behave like an identity function, passing information without meaningful transformation.
- Pre-Layer Normalization (Pre-LN)
- A training approach designed to stabilize model training. However, it is susceptible to variance exploding exponentially with depth, which leads to the deep layers becoming redundant and less effective for computation.
- LayerNorm Scaling (LNS)
- The proposed solution that mitigates exponential variance growth by scaling the layer normalization output by 1/sqrt(ℓ), where ℓ is the depth. This simple modification stabilizes training and enhances performance.
Terminology used across episodes
This episode discusses
- The Curse of Depth in Large Language Models · Paper Radio
- GPT-4 Technical Report
- Layer Normalization
- Qwen2.5-VL Technical Report
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- GLM: General Language Model Pretraining with Autoregressive Blank Infilling
- Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
- OLMo: Accelerating the Science of Language Models
- The Unreasonable Ineffectiveness of the Deeper Layers
- GradientStabilizer:Fix the Norm, Not the Gradient
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
- Mistral 7B
- The Remarkable Robustness of LLMs: Stages of Inference?
- Outlier-weighed Layerwise Sampling for LLM Fine-tuning
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
- OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- GLU Variants Improve Transformer
- Spike No More: Stabilizing the Pre-training of Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
The paper
The Curse of Depth in Large Language Models · Read on arXiv
Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu
Westlake University, China · Emory University, USA · Dalian University of Technology, China · University of Surrey, UK · University of Oxford, UK
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Curse of Depth in Large Language Models".
Jane: The paper was written by Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng et al. from Westlake University, China and Emory University, USA and Dalian University of Technology, China and University of Surrey, UK and University of Oxford, UK.
Tom: Stay tuned as we take you through the paper and discuss its implications.
The Core Problem: Tom: So, "The Curse of Depth in Large Language Models" confirms that this problem isn's just an isolated bug; the phenomenon of ineffective deep layers exists across major LLM families like Llama and DeepSeek. This is a widespread issue, not a niche problem.
Jane: And the paper provides some really clear evidence, showing things like performance drop tests where removing one layer after its depth shows very little change in model performance at all levels. It’s shocking how resilient those deeper layers are.
Meng: If pruning deep layers barely hurts the overall accuracy, it strongly suggests that these layers aren't contributing much to the final output, which is a massive practical hurdle for optimization algorithms.
Lu: The paper's analysis points to a specific mechanism: Pre-Layer Normalization (Pre-LN). This approach is designed to stabilize training but doesn'output variance explodes exponentially with depth.
Lalam: That exponential growth of variance, as the model gets deeper, means the layers are essentially collapsing into an identity function. They aren't performing meaningful transformations anymore; they are just passing information along without changing it much.
Tom: The paper proves that when this happens, the derivative of those deep blocks approaches identity matrix. It’s mathematically proven that's why they become so ineffective.
Meng: This is where the efficiency problem really hits the fan, Tom. If we are paying for layers to behave like a simple pass-through function, we have to rethink how these massive architectures are constructed entirely differently.
The Solution: Tom: Thankfully, "The Curse of Depth in Large Language Models" suggests a clean fix: LayerNorm Scaling, or LNS. This is the proposed solution to mitigate the exponential variance growth.
Jane: LNS scales the output of layer normalization by /sqrt based on its depth. It’s a simple modification, but it’s incredibly powerful because it actively controls that variance explosion before it can happen.
Meng: I appreciate that LNS is hyperparameter-free and doesn't introduce extra parameters during training. That makes implementation extremely straightforward for large-scale deployment.
Lu: The theoretical implications are huge, Lu thinks, because by slowing the growth of variance from exponential to at most quadratic, we’re fundamentally changing the landscape of what a deep neural network can achieve. It expands our theoretical capacity dramatically.
Lalam: LNS helps ensure that every single layer contributes something unique and different to the final representation. Instead of just stacking redundant functions, we get a truly diverse set of features from every layer in the model.
Tom: Exactly, Lalam; it makes sure those deep layers are contributing meaningfully again. This approach is designed to counteract that identity mapping behavior and revitalize the contribution of those later blocks.
The Results: Tom: We’ve seen how the problem manifests, but now we’re looking at the empirical evidence in "The Curse of Depth in Large Language Models." The experiments show LNS consistently outperforms all other existing normalization techniques.
Jane: It shows a substantial improvement in both pre-training performance and when we move to supervised fine-tuning tasks. The model is learning better representations across different sizes, from 130M up to 7B parameters.
Meng: And the results are scalable; LNS works whether you're training a small Llama variant or a massive state-of-the-art OLMo model. It consistently delivers lower perplexity and more stable training dynamics for large systems.
Lu: The fact that the angular distance between layers increases under LNS also confirms that the representation is truly diversifying, not just because the variance is controlled, but because of how the signals are propagating through depth.
Lalam: This means we can finally build AI systems where every single step in every layer matters. It's a huge win for resource efficiency and quality simultaneously.
Tom: It’s clear that by fixing the mechanism of Pre-LN, we’ve successfully tackled the Curse of Depth and found a robust alternative.
Wrap Up: Tom: As we wrap up our discussion on "The Curse of Depth in Large Language Models," it seems like we've seen a major breakthrough in how these massive models are built and trained.
Jane: It’s genuinely exciting that a simple scaling factor, based on the square root of depth, can unlock so much more performance potential across different models.
Lu: I think the biggest intellectual takeaway is that we have successfully moved away from an exponential constraint toward a manageable polynomial growth in variance. That's a huge shift in our understanding of deep learning theory.
Meng: From an engineering viewpoint, this means that we can move forward with these architectures knowing that the performance gains aren't just theoretical; they are practical and scalable across massive training budgets.
Lalam: I hope this work encourages more leads to build truly efficient AI systems for the future, ensuring that every computational step serves a meaningful purpose.
Tom: Absolutely, Lalam. It’s clear that "The Curse of Depth in Large Language Models" offers a robust and computationally efficient path forward for the next generation of AI. Goodbye everyone, and we'll see you on the next paper!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization