Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters
cs.LG, cs.CL
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/kunikrubika05/step-law-small-scale
Terminology
Sources
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Scaling Laws for Neural Language Models
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Adam: A Method for Stochastic Optimization
- Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
- On the Variance of the Adaptive Learning Rate and Beyond
- Decoupled Weight Decay Regularization
- An Empirical Model of Large-Batch Training
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks