Toward a First-Principles Update Geometry for the Language-Model Head
cs.LG
Submitted: 2026-08-23
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by/4.0/
The gist: Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers.
Terminology
Abstract
Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional change. Softmax removes shared logit shifts, whereas the spectral norm can assign arbitrarily large size to updates that change no output probability. We therefore treat the LM head and softmax as one module and derive an update geometry for their composition. Hilbert's projective distance respects this invariance as it measures the largest change in pairwise log odds. For an update S with token rows s i, we show that the largest Hilbert distance over h 2 at most H is exactly H D(S), where D(S)= i<j s i - s j 2 is the Euclidean row diameter. This diameter replaces the spectral norm in the resulting Muon-style steepest descent problem. An exact solution is possible, but its direct formulation contains one d-dimensional vector variable for every token pair. For a vocabulary size of approximately 50 k, this means more than one billion token pairs, making the calculation impractical at every training step. We instead impose a stronger common-ball constraint and derive projected RowNorm as an O(Vd) solution. For the exact RowNorm oracle, we prove that its first-order decrease is at least 1/sqrt 2 of the exact diameter-constrained optimum. With Muon on the backbone, experiments across three seeds at 190M, 380M, and 640M parameters show that RowNorm reduces mean final step diameters and empirical Hilbert RMS perturbations by factors of 45 -- 60 and 12 -- 15, respectively, with only a 0.0057 -- 0.0153 increase in mean final validation loss.
Sources
- Modular Duality in Deep Learning
- Minimal $N$-Point Diameters and $f$-Best-Packing Constants in $R^d$
- Hyperbolic contractivity and the Hilbert metric on probability measures
- On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- Generalized Neural Collapse for a Large Number of Classes
- Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
- Muon is Scalable for LLM Training
- Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
- Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks