Optimal Transport Depth Up-Scaling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Optimal Transport Depth Up-Scaling".
Jane: The gist Scaling Large Language Models (LLMs) yields performance gains but incurs substantial training costs,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We've talked about how this paper calls Optimal Transport Depth Up-Scaling, which is really about tackling the performance gains from scaling up language models while keeping training costs manageable.
Jane: The main thesis of this work is that existing methods for depth up-scaling often fall short because they don't account for the way neurons are permuted across layers, which can cause misalignment and hurt the final model quality.
Lu: They propose OpT-DeUS to fix this by using Optimal Transport to align and fuse adjacent Transformer blocks when creating new layers, ensuring a better neuron correspondence between those layers.
Meng: It’s essentially about finding a mathematically optimal way to initialize these new blocks so they respect the underlying structure of the base model during the scaling process.
Tom: They claim that this method achieves better overall performance and improved training efficiency when comparing continual pre-training and supervised fine-tuning across different expanded model sizes.
Jane: And they found that by focusing on where you insert those new layers—specifically near the top—you get higher training efficiency because it shortens the backpropagation paths for those newly added layers.
Lu: The paper is important because it shows a concrete way to integrate structural alignment, using concepts from Optimal Transport, into model expansion techniques to get better results without just relying on simple weight copying.
Meng: From an engineering view, this gives us a more robust initialization strategy that should lead to higher quality models for the same training budget.
Conclusion: Tom: So, to wrap up, Optimal Transport Depth Up-Scaling is a method that uses math to align layers using Optimal Transport principles instead of just copying weights.
Jane: The authors did a lot of work showing that this approach results in better performance and better speed during training compared to other current depth up-scaling techniques.
Lu: They proved that inserting new layers near the top leads to faster training because it shortens the backpropagation path for those new layers, which is a key finding.
Meng: This implies that for someone building large language models, you should consider how you structure your scaling process to maximize efficiency by placing those new layers strategically.
Tom: So what this means is that this paper gives us a concrete tool—OpT-DeUS—to make expanding LLMs more stable and efficient by paying attention to the internal connections between the layers.
Jane: It’s about moving beyond simple weight copying toward a method that respects the functional structure of the model while scaling up, which is what this work does.
Lu: The implication is that future research might explore how to combine these transport matrix concepts with even more complex structural changes in model architecture to see what else we can achieve.
University of Sheffield
cs.CL
Submitted: 2025-08-11
Updated: 2026-10-07
Code: https://github.com/voalmciaf/OpT-DeUS
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 91/100
The gist: The gist Scaling Large Language Models (LLMs) yields performance gains but incurs substantial training costs, and this paper proposes Optimal Transport Depth Up-Scaling (OpT-DeUS) to mitigate neuron
Key concepts
- Neuron Permutation Mismatch
- When adding new layers to a model, the neurons in the new layer might not correspond correctly to the neurons in the base layers. Copying or averaging weights directly can break this correspondence, leading to performance degradation.
- Optimal Transport (OT)
- OT is a mathematical tool used here to find an optimal way to move or align one set of data points (neurons from one layer) onto another set (neurons in the next layer). This alignment helps create new layers with neuron-aligned structures, preventing mismatches.
- Progressive Interpolation
- This is a method where new layers are added incrementally to an existing model. OpT-DeUS specifically uses this by updating only the newly added layers at each step, using OT for alignment to ensure smooth and effective scaling.
Terminology
Summary
The gist Scaling Large Language Models (LLMs) yields performance gains but incurs substantial training costs, and this paper proposes Optimal Transport Depth Up-Scaling (OpT-DeUS) to mitigate neuron permutation mismatch between layers by aligning and fusing Transformer blocks using Optimal Transport for new layer creation.
Motivation
The motivation behind OpT-DeUS stems from the fact that existing methods often copy or average weights from base layers, which neglects neuron permutation differences that can harm performance. Same-indexed neurons from different layers may not be functionally corresponding, directly copying or averaging them can harm downstream performance. This limitation motivates the main research question: How to effectively initialize new layers to avoid neuron permutation mismatches in progressive depth up-scaling. Inspired by applying Optimal Transport (OT), OpT-DeUS aligns and fuses adjacent layers block-wise to create neuron-aligned new layers.
How it works
OpT-DeUS is a progressive interpolation method that updates only the newly added layers. It aligns and fuses adjacent base layers via OT for new layer creation. The process involves several steps detailed in Algorithm 1.
The weight initialization process consists of five steps detailed in Algorithm 1.
Step-2 involves alignment within the layer where the permutation change caused by aligning W(i)b-1 to W(i)b disrupts the original neuron correspondence between W(i)b-1 and W(i)b. This is restored by performing alignment via Tin, which is defined by TMF for each b in f′.
Step-3 involves alignment across the layer where we solve OT(α, β, C) to compute the transport matrix T that minimizes Pk,j Tk ckj subject to the marginal constraints T1m = α and TT 1n = β. This is employed to obtain Tout for W(i)b.
Step-4 computes f′i weights where W′(i)b is initialized by averaging the aligned W(i)b and W(i+1)b.
Step-5 involves zero-initialization where WO = 0 and Wdown = 0, which naturally resolves misalignment issues while ensuring function preservation.
Experimental Setup
The experiments utilize the 32-layer Llama-3.1-8B (Grattafiori et al. 2024) as the base model for the 11.5B expanded models. The expanded model sizes are fixed at 11.5B parameters with 48 layers (adding 16 layers) and 1.72B with 24 layers (adding 8 layers) for all depth up-scaling methods.
The evaluation includes both continual pre-training (CPT) and supervised fine-tuning (SFT) stages. For CPT, the data used is sampled from the CC-MAIN-2024-51 subset of FineWeb-Edu. For SFT, Alpaca GPT4 (Peng et al. 2023) is chosen and the whole model is updated.
Results and Analysis
OpT-DeUS achieves top performance on five out of eight benchmarks during CPT and SFT for the 11.5B expanded models. For the 1.72B expanded models, OpT-DeUS achieves the best overall performance (52.02) and ranks first on Wiki-PPL (12.19), LogiQA (22.58), and CSQA (43.00).
An ablation study on interpolation positions shows that OpT-DeUS-Top is the best performing strategy overall, yielding the highest average performance (68.87). This suggests that inserting new layers closer to the top results in higher training efficiency due to shorter backpropagation time while obtaining additional performance gains.
OpT-DeUS consistently outperforms Avg-DeUS on both 11.5B and 1.72B expanded models (Avg: 68.87 vs 68.39; 52.02 vs 50.88). This consistent improvement confirms that using OT for neuron alignment during initialization comprehensively enhances the downstream performance of progressive depth up-scaling.
The analysis on training efficiency reveals a strong correlation between interpolation positions and efficiency: tophalf insertions are notably faster, while bottom half insertions require longer training time. OpT-DeUS-Top (12:52:04) is notably faster than OpT-DeUSBtm (14:56:00).
The paper concludes that OpT-DeUS offers better downstream performance with improved training efficiency than other depth up-scaling approaches. The analysis of interpolation positions reveals their impact on training efficiency, demonstrating that inserting new layers closer to the top leads to higher training efficiency due to shorter backpropagation paths through the trainable new layers.
The paper also notes that LESA’s perplexity sharply increases when applied to Llama-3.2-1B (871.50), suggesting that smaller models have fewer layers, leading to less training data for the auxiliary network, consequently causing it to underfit. The paper hypothesizes that SOLAR’s poor performance is caused by catastrophic forgetting because fully updating the expanded model substantially degrades the pre-trained parametric knowledge. The paper also observes that OpT-DeUS achieves top training efficiency among baselines.
The paper compares LESA and OpT-DeUS in terms of creation and training time, noting that OpT-DeUS requires less time compared to LESA. Both methods require additional computation, but OpT-DeUS achieves the best time efficiency among the baselines. The paper hypothesizes that this increased time for LESA is mainly caused by the extra computation required for SVD when scaling up base models.
The paper concludes that OpT-DeUS offers better downstream performance with improved training efficiency than other depth up-scaling approaches. The analysis of interpolation positions reveals their impact on training efficiency, demonstrating that inserting new layers closer to the top leads to higher training efficiency due to shorter backpropagation paths through the trainable new layers. The paper also observes that OpT-DeUS consistently achieves top performance on at least four out of eight benchmarks across all checkpoints regardless the size of the CPT data. The paper further notes that both LLaMA-Pro and OpT-DeUS match the base model’s perplexity regardless of model parameters due to function preservation, demonstrating maximum expansion stability compared to other baselines.
The paper also observes that OpT-DeUS consistently achieves top performance on at least four out of eight benchmarks across all checkpoints regardless the size of the CPT data. The paper further notes that both LLaMA-Pro and OpT-DeUS match the base model’s perplexity regardless of model parameters due to function preservation, demonstrating maximum expansion stability compared to other baselines.
The paper also observes that OpT-DeUS consistently achieves top performance on at least four out of eight benchmarks across all checkpoints regardless the size of the CPT data.
Improvements for AI systems
- Bold header: Optimal Transport Depth Up-Scaling (OpT-DeUS) for Neuron Alignment
This method aligns and fuses adjacent Transformer blocks via OT for new layer creation, to mitigate neuron permutation mismatch between layers,
which directly addresses the limitation of existing methods that copy or average weights from base layers, neglecting neuron permutation differences.
- Bold header: Improved Training Efficiency via Interpolation Position Analysis
The analysis shows that inserting new layers closer to the top results in higher training efficiency due to shorter back-propagation time while obtaining additional performance gains,
suggesting a strategic insertion policy for model expansion.
- Bold header: Enhanced Downstream Performance Across Model Scales
OpT-DeUS achieves top performance on five out of eight benchmarks
for both Continual Pre-training and Supervised Fine-Tuning across different model sizes, demonstrating robustness to model sizes
compared to baselines like SOLAR, which showed poor performance.
Sources
- Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Llama 3 Herd of Models
- Kuwain 1.5B: An Arabic SLM via Language Injection
- Scaling Laws for Neural Language Models
- Counting Carbon: A Survey of Factors Influencing the Emissions of Machine Learning
- Instruction Tuning with GPT-4
- Progressive Neural Networks
- The Computational Limits of Deep Learning
- Will we run out of data? Limits of LLM scaling based on human-generated data
- Layers at Similar Depths Generate Similar Activations Across LLM Architectures
- Progressively Stacking 2.0: A Multi-stage Layerwise Training Method for BERT Training Speedup
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering