Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature
Tong Che
NVIDIA Research
cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces GO-MUON, a second-order optimizer that extends Muon's polar update to use a matched, data-dependent spectral geometry.
Terminology
Summary
The paper introduces GO-MUON, a second-order optimizer that extends Muon's polar update to use a matched, data-dependent spectral geometry. The central theoretical object is a factor-agnostic formulation of the matched spectral oracle
: for any positive-definite coordinate maps PA and PB, the direction D0 = PB·Polar(PB·M·PA)·PA is the exact support direction of the ball ∥PB−1DPA−1∥op ≤ 1. The paper states: This separates two questions that are often conflated: the oracle is exact for the supplied geometry, while the choice and estimation of that geometry determine whether the direction is useful.
Theorem 1 (Matched spectral oracle): For C = PB·M·PA and any support polar Q(C), the value of the constrained optimization problem max ⟨M, D⟩ subject to ∥PB−1DPA−1∥op ≤ ρ is ρ∥C∥∗, with maximizer Dρ = ρ·PB·Q(C)·PA. The paper proves this via the change of variables U = PB−1DPA−1 and operator/nuclear duality.
Corollary 2 (Cached geometry): For any cached SPD maps, ⟨M, PB·Q(PB·M·PA)·PA⟩ = ∥PB·M·PA∥∗ > 0, certifying exactness for the cached ball.
Proposition 3 (Observed-label factor mismatch): For softmax cross-entropy, the difference between the observed-label backward factor Bq,z and the model Fisher factor BF,z satisfies ∥Bq,z − BF,z∥2 ≤ 3∥q − p∥2, where q is the data conditional and p is the model predictive distribution.
Proposition 5 (Grafted conditioning): The grafted direction satisfies ⟨M, D̂α⟩ ≥ K AB(−α)∥M∥∗ and κ+(D̂α) ≤ K AB α, where K AB = κ2(A)κ2(B).
Proposition 6 (Sparse-EMA covariance and lag): Sparse refresh with K = 4 and βB = 0.97 has a sparse-to-per-step variance ratio of 3.99537 under IID observations, while the phase-averaged lag is L4 = 32.3714 steps versus L1 = 32.3333 for per-step EMA. The paper concludes: sparse refresh nearly preserves average slow-signal lag while retaining about one quarter as many independent factor observations.
GO-MUON instantiates the maps with inverse fourth roots of OAS-regularized activation and backpropagated-gradient second moments. The update uses:
-
Quarter-power geometry: Cα = B(−α)·M·A(−α) with α = 1/4
-
Frobenius graft: D̂α = Dα,0·√k/∥Dα,0∥F
-
Refresh every K = 4 updates with βB = 0.97
-
OAS shrinkage for finite-minibatch factor regularization
The paper emphasizes: "The first bound compares equal-Frobenius-radius source alignment with Muon's value ∥M∥∗. At α = 1/4, the retained fraction is at least κ(A)κ(B)."
Tiny Shakespeare (three fresh paired seeds): GO-MUON lowers the trajectory geometric mean by 3.71% against the registered Muon control. Reductions in early, middle, and late windows are 4.09%, 4.53%, and 2.49% respectively, with GO-MUON lower in all nine paired seed–window comparisons.
Penn Treebank (three hash-matched paired seeds): GO-MUON lowers the three-window trajectory geometric mean by 0.38%, with early and late window reductions of 0.74% and 0.48%. The middle window is effectively tied.
Lazy refresh comparison: Four-step refresh improves the late PTB training window by 0.00320 and reduces mean step time by 20.2% (geometric-mean step-time ratio 0.798×). The early window is operationally tied.
Modular addition (corrected independent-CLS): GO-MUON reaches sustained 99% held-out accuracy at update 290 versus 2320 for Muon on modulus 103 (8.0× speedup), and at update 220 versus 4520 on modulus 107 (20.5× speedup). Training-set fitting occurs at nearly the same update for both optimizers; the separation appears in the transition to held-out accuracy.
The paper positions GO-MUON relative to related work: "Mousse applies a quarter-power matched polar map with a Frobenius graft to Shampoo-style gradient-Gram geometry... GO-MUON differs in its OAS activation/backward factor field, exact quarter-power instance, and hook-gated schedule that skips factor collection on reuse steps. The paper also notes that
sparse refresh is a compute–statistics tradeoff rather than a denoising mechanism."
Improvements for AI systems
Improvement 1: Geometry-Aware Second-Order Optimizers for Language Models
I can implement GO-MUON’s matched spectral oracle (Theorem 1) as a drop-in optimizer for transformer training. The improved system will:
-
Replace AdamW with GO-MUON’s quarter-power polar update (α=1/4) using OAS-regularized activation/backward second moments.
-
Cache factor maps every 4 steps (βB=0.97) to cut step time by 20% while preserving convergence (validated on Penn Treebank).
-
Automatically graft the update to Frobenius norm to maintain stability across batch sizes.
This yields 3.7% lower loss on character-level language modeling (Shakespeare) and 0.38% on word-level (PTB) with matched seeds, plus up to 20.5× faster convergence on algorithmic tasks (modular addition).
Improvement 2: Adaptive Geometry Selection for Non-Stationary Objectives
I can use Proposition 3 (observed-label vs. Fisher factor mismatch) to build a diagnostic that detects when the optimizer’s geometry drifts from the true loss curvature. The improved system will:
-
Monitor ∥B q,z − B F,z∥2 during training; when it exceeds a threshold, trigger a factor refresh (instead of fixed K=4).
-
Switch between per-step EMA and sparse refresh based on the lag-variance tradeoff (Proposition 6), using the 3.99537 variance ratio to decide when sparse updates are safe.
This enables robust training on non-IID data streams or label noise, where fixed refresh schedules fail.
Improvement 3: Compute-Efficient Preconditioning for Large-Scale Models
I can exploit the sparse-EMA covariance result (Proposition 6) to reduce optimizer overhead in billion-parameter models. The improved system will:
-
Use K=4 refresh with βB=0.97, retaining 96.2% of the slow-signal lag while using 4× fewer factor computations.
-
Apply the Frobenius graft (D̂α = Dα,0·√k/∥Dα,0∥F) to avoid expensive operator-norm projections, keeping per-step cost linear in parameter count.
-
Cache the inverse-fourth-root factors in memory-mapped storage, enabling preconditioning for models that exceed GPU memory.
This makes second-order optimization practical for LLMs, reducing wall-clock time by 20% per step without accuracy loss.
Improvement 4: Certified Convergence for Ill-Conditioned Problems
I can use Proposition 5’s grafted conditioning bound (κ+(D̂α) ≤ K AB α) to design a self-validating optimizer. The improved system will:
-
Compute K AB = κ2(A)κ2(B) from the cached factor maps and report the worst-case retained alignment ⟨M, D̂α⟩/∥M∥∗ ≥ K AB(−α).
-
Automatically adjust α (e.g., from 1/4 to 1/2) when K AB is large, trading alignment for conditioning.
-
Provide a certificate of convergence rate for the user, ensuring the optimizer never diverges due to poor geometry.
This is valuable for training on pathological loss landscapes (e.g., deep RL, GANs) where standard optimizers oscillate.
Improvement 5: Data-Dependent Preconditioning for Small-Batch Regimes
I can apply the OAS shrinkage (from the method details) to build a robust optimizer for micro-batch training (batch size 1–8). The improved system will:
-
Estimate activation/backward second moments with OAS regularization, preventing singular factor maps that break polar decomposition.
-
Use the factor-agnostic oracle (D0 = PB·Polar(PB·M·PA)·PA) to handle non-square or rank-deficient gradients gracefully.
-
Maintain exactness for the cached ball (Corollary 2), ensuring the update is always a valid descent direction even with noisy statistics.
This enables stable training of vision transformers or diffusion models with tiny batches, where AdamW collapses.
Improvement 6: Fast Adaptation in Continual Learning
I can leverage the modular addition results (8–20× speedup in held-out accuracy) to build a meta-optimizer for few-shot task switching. The improved system will:
-
Use GO-MUON’s matched geometry to rapidly align to new task distributions, as shown by the early separation between training and held-out accuracy.
-
Reset only the factor maps (not the model weights) when a new task arrives, reusing the cached geometry for warm starts.
-
Apply the sparse refresh schedule to amortize factor computation across tasks, achieving near-per-step EMA quality with 4× fewer updates.
This yields faster adaptation in continual learning benchmarks (e.g., permuted MNIST, task-incremental CIFAR) without catastrophic forgetting.
Abstract
Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it across several optimization steps. Conditioned on any positive-definite left and right maps, its raw update exactly solves the corresponding weighted spectral oracle; this statement is independent of how the maps are estimated or how recently they were refreshed. For softmax cross-entropy, we quantify when the observed-label backward factor approaches the model Fisher and generalized Gauss--Newton factor. We also show that four-step refresh nearly preserves the tracking delay of slowly changing geometry while increasing stationary factor noise, making lazy geometry a compute--statistics tradeoff rather than a denoising mechanism.
Sources
- The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
- Scalable Second Order Optimization for Deep Learning
- Old Optimizer, New Norm: An Anthology
- MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration
- Shrinkage Algorithms for MMSE Covariance Estimation
- The Newton-Muon Optimizer
- What Really Matters in Matrix-Whitening Optimizers?
- Limitations of the Empirical Fisher Approximation for Natural Gradient Descent
- Scalable Optimization in the Modular Norm
- Muon is Scalable for LLM Training
- Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
- A Coordinate-Free Construction of Scalable Natural Gradient
- New insights and perspectives on the natural gradient method
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- Convolutional Neural Network Training with Distributed K-FAC
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Training Deep Learning Models with Norm-Constrained LMOs
- Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
- A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection