Optimal Learning Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
stat.ML, cs.LG
Submitted: 2026-02-06
Updated: 2026-09-14
Comments: Accepted at COLT 2026. Major revision with improved analysis of fractional LR schedule
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training dynamics into signal learning and noise forgetting.
Terminology
Abstract
We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training dynamics into signal learning and noise forgetting. In power-law kernel regression, these two components are governed by a source exponent s>0 and a capacity exponent q>1, respectively, with smaller s corresponding to harder tasks. For a fixed training horizon N, we characterize the schedules that minimize the final-step loss under a stability constraint and reveal a sharp phase transition. In the easy-task regime s>1-1/q, the optimal schedule follows power decay from the beginning of training; in the hard-task regime s<1-1/q, it becomes warmup-stable-decay (WSD)-like (Hu et al., 2024), staying at the largest admissible LR for most of training before a final decay. In both regimes, the decay exponent is 2q-1: task difficulty determines when to decay, while model capacity determines how to decay. Beyond the exact optimum, we study fractional schedules, whose shape is defined over relative training progress. We show that precise tuning of the decay shape is often unnecessary: a broad class of profiles attains the optimal convergence rate, while overly slow terminal decay leads to schedule-induced capacity saturation. Finally, for one-pass SGD in kernel regression, FSL-motivated power-decay schedules achieve optimal last-iterate rates. Experiments support the theoretical predictions and the task-dependent transition between early and delayed decay.
Sources
- Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
- A Dynamical Model of Neural Scaling Laws
- Convex Optimization: Algorithms and Complexity
- Optimal Linear Decay Learning Rate Schedules and Further Refinements
- Scaling Laws for Autoregressive Generative Modeling
- Deep Learning Scaling is Predictable, Empirically
- Training Compute-Optimal Large Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Scaling Laws for Hyperparameter Optimization
- Scaling Laws for Neural Language Models
- Scaling Laws for Precision
- Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
- A simpler approach to obtaining an O(1/t) convergence rate for the projected stochastic subgradient method
- Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
- Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
- Improved Scaling Laws in Linear Regression via Data Reuse
- Scaling Laws in Linear Regression: Compute, Parameters, and Data
- DeepSeek-V3 Technical Report
- SGDR: Stochastic Gradient Descent with Warm Restarts
- A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey