On the origin of neural scaling laws: from random graphs to natural language
cs.LG, cond-mat.dis-nn, cs.AI, stat.ML
Submitted: 2026-01-15
Updated: 2026-01-15
Code: https://github.com/dayal-kalra/nanoGPT-mup
Terminology
Sources
- Explaining Neural Scaling Laws
- Chinchilla Scaling: A replication attempt
- Establishing Task Scaling Laws via Compute-Efficient Model Ladders
- PaLM: Scaling Language Modeling with Pathways
- Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
- Language models scale reliably with over-training and on downstream tasks
- The Llama 3 Herd of Models
- Scaling Laws for Autoregressive Generative Modeling
- Deep Learning Scaling is Predictable, Empirically
- Training Compute-Optimal Large Language Models
- Learning Curve Theory
- Mixtral of Experts
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Scaling Laws for Neural Language Models
- Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Superposition Yields Robust Neural Scaling
- A Solvable Model of Neural Scaling Laws
- The Quantization Model of Neural Scaling
- Scaling Data-Constrained Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks