Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
cs.LG, math.OC
Submitted: 2025-09-15
Updated: 2026-09-17
Comments: 26 pages, add numerical comparison with Galore and SOAP
Code: https://github.com/dengzhanwang/Low-rank-Muon
Project page: http://leloykun.github.io/ponder/steepest-descent-non-riemannian
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlooked.
Terminology
Abstract
Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlooked. Recently, the optimizer Muon, which explicitly exploits this structure, has gained significant attention for its strong performance in foundation model training. A key component contributing to Muon's success is matrix orthogonalization. In this paper, we propose low-rank orthogonalization, which performs orthogonalization by leveraging the low-rank nature of gradients during NN training. Building on this, we introduce low-rank matrix-signed gradient descent (MSGD) and a low-rank variant of Muon. %Numerical experiments demonstrate the superior performance of low-rank orthogonalization, with low-rank Muon achieving promising results in GPT-2 and LLaMA pretraining---surpassing the carefully tuned vanilla Muon on tasks with large model sizes. Numerical experiments demonstrate the advantages of low-rank orthogonalization: low-rank Muon generally matches or improves upon vanilla Muon on the GPT-2 and LLaMA pretraining tasks, with clearer improvements observed for relatively larger models. Theoretically, we establish the iteration complexity of low-rank MSGD for finding an approximate stationary solution, and the iteration complexity of low-rank Muon for finding an approximate stochastic stationary solution under heavy-tailed noise. The code to reproduce our numerical experiments is available at https://github.com/dengzhanwang/Low-rank-Muon.
Sources
- Dion: Distributed Orthonormalized Updates
- Practical Efficiency of Muon for Pretraining
- ASGO: Adaptive Structured Gradient Optimization
- Scalable Second Order Optimization for Deep Learning
- Muon Optimizes Under Spectral Norm Constraints
- Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
- Complexity of normalized stochastic first-order methods with momentum under heavy-tailed noise
- Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization
- PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
- COSMOS: A Hybrid Adaptive Optimizer for Memory-Efficient Training of LLMs
- Training Deep Learning Models with Norm-Constrained LMOs
- Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
- Convergence Bound and Critical Batch Size of Muon Optimizer
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- On the Convergence Analysis of Muon
- AdaMuon: Adaptive Muon Optimizer
- LLaMA: Open and Efficient Foundation Language Models
- Orthogonalising gradients to speed up neural network optimisation
- Structured Preconditioners in Adaptive Optimization: A Unified Analysis
- ADADELTA: An Adaptive Learning Rate Method
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks