Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/KellerJordan/Muon
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) pretraining conventionally returns the raw final iterate.
Terminology
Abstract
Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose Terminal Shrinkage Averaging (TSA), which interpolates between the raw final iterate and the average of recent checkpoints to balance recent progress against terminal variation. We analyze how TSA changes the preferred terminal learning-rate schedule under a local quadratic approximation and test this interaction through a sequence of controlled NanoChat experiments. Finally, we demonstrate that the resulting gains transfer to depth-22 NanoChat, where the combined schedule and estimator improve validation quality. A qualifying time-to-GPT-2 run also finishes faster than the public baseline used in our experiments, providing preliminary evidence of benchmark acceleration.
Sources
- Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models
- The Road Less Scheduled
- Training Compute-Optimal Large Language Models
- Stop Wasting My Time! Saving Days of ImageNet and BERT Training with Latest Weight Averaging
- Scaling Laws for Neural Language Models
- DataComp-LM: In search of the next generation of training sets for language models
- Muon is Scalable for LLM Training
- Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
- Iterate averaging as regularization for stochastic gradient descent
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- SOAP: Improving and Stabilizing Shampoo using Adam
- Fantastic Pretraining Optimizers and Where to Find Them
- Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks