Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
cs.LG
Submitted: 2026-03-17
Updated: 2026-09-27
Terminology
Sources
- Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
- Effective Quantization of Muon Optimizer States
- Adam: A Method for Stochastic Optimization
- Muon is Scalable for LLM Training
- FP8 Formats for Deep Learning
- A Theory on Adam Instability in Large-Scale Machine Learning
- FP8-LM: Training FP8 Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks