Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
stat.ML, cs.LG, math.OC
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
- Logarithmic-time Schedules for Scaling Language Models with Momentum
- The Llama 3 Herd of Models
- Deep Learning Scaling is Predictable, Empirically
- Learning Curve Theory
- Scaling Laws for Neural Language Models
- Optimal Learning Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
- Accelerating SGD with momentum for over-parameterized learning
- Muon is Scalable for LLM Training
- A Solvable Model of Neural Scaling Laws
- Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods
- 4+3 Phases of Compute-Optimal Neural Scaling Laws
- Benchmarking Optimizers for Large Language Model Pretraining
- Practical Efficiency of Muon for Pretraining
- Asymmetric Scaling Laws from Sparse Features
- Convergence of Momentum-Based Optimization Algorithms with Time-Varying Parameters
- The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
- A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
- The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey