Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
cs.LG, cs.AI, stat.ML
Submitted: 2026-05-26
Updated: 2026-08-28
Code: https://github.com/karpathy/nanoGPT
Terminology
Sources
- Layer Normalization
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
- Training Compute-Optimal Large Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Adam: A Method for Stochastic Optimization
- Muon is Scalable for LLM Training
- Decoupled Weight Decay Regularization
- Power-law escape rate of SGD
- 2 OLMo 2 Furious
- A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
- Gemma 2: Improving Open Language Models at a Practical Size
- LLaMA: Open and Efficient Foundation Language Models
- A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
- Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis
- Qwen2 Technical Report
- Qwen3 Technical Report
- Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
- HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks