Beyond the Matrix Sign: Quadratic Spectral Descent
cs.LG, cs.AI
Submitted: 2026-09-07
Updated: 2026-09-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball.
Terminology
Abstract
Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. This gives a matrix-sign update that preserves the singular directions of the gradient and assigns the same magnitude to all active singular modes. We ask whether these two properties remain optimal when local curvature is taken into account. To answer this question, we keep Muon's spectral-norm constraint unchanged and replace the linear local model with a quadratic one. We call the resulting method Quadratic Spectral Descent (QSD). We show that curvature can change both the singular values and the singular directions of the optimal update. To make QSD practical, we approximate curvature with Kronecker-factored statistics and solve the constrained quadratic with a small number of Frank--Wolfe steps, each of which has a closed-form matrix-sign subproblem. We further provide an optimality certificate, a comparison with Muon under the same quadratic surrogate, and an O(1/K) convergence rate for the inner solver. Experiments on GPT pre-training show that QSD consistently improves validation loss over Muon and recent Muon variants, and reduces wall-clock training time by up to 8.49% at matched validation loss.
Sources
- Old Optimizer, New Norm: An Anthology
- Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
- Muon$^p$: Muon with Fractional Spectral Powers
- The Newton-Muon Optimizer
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- NorMuon: Making Muon more efficient and scalable
- Muon is Scalable for LLM Training
- Decoupled Weight Decay Regularization
- Training Deep Learning Models with Norm-Constrained LMOs
- Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
- Practical Efficiency of Muon for Pretraining
- Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- Why Muon Outperforms Adam: A Curvature Perspective
- DynMuon: A Dynamic Spectral Shaping View of Muon
- Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
- FISMO: Fisher-Structured Momentum-Orthogonalized Optimizer
- Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks