Convergence guarantees for Muon: New parameter regimes and generalizations
math.NA, cs.AI, cs.NA, math.OC
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/kellerjordan/modded-nanogpt
Terminology
Sources
- The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
- Modular Duality in Deep Learning
- On the Convergence of Muon and Beyond
- Convergence of Muon with Newton-Schulz
- A Note on the Convergence of Muon
- Muon is Scalable for LLM Training
- Convergence Bound and Critical Batch Size of Muon Optimizer
- On the Convergence Analysis of Muon
Related papers
- Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm
- A Neural-preconditioned Poisson Solver for Mixed Dirichlet and Neumann Boundary Conditions
- Second-order consistency for learning chaotic dynamics via randomized Jacobian matching
- Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
- Data-efficient Kernel Methods for Learning Hamiltonian Systems
- Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems