MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
cs.LG, cs.CL
Submitted: 2026-05-26
Updated: 2026-09-22
Comments: Findings of the Association for Computational Linguistics: EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates.
Terminology
Abstract
The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term introduces curvature-sensitive corrections while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.
Sources
- Dion: Distributed Orthonormalized Updates
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
- FullStack Bench: Evaluating LLMs as Full Stack Coders
- Training Verifiers to Solve Math Word Problems
- Drop-Muon: Update Less, Converge Faster
- MuonBP: Faster Muon via Block-Periodic Orthogonalization
- CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
- Adam: A Method for Stochastic Optimization
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Muon is Scalable for LLM Training
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- Scaling Embeddings Outperforms Scaling Experts in Language Models
- Evaluating Language Models for Efficient Code Generation
- Decoupled Weight Decay Regularization
- Mixed Precision Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks