OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling
cs.LG, cs.CL
Submitted: 2026-05-08
Updated: 2026-08-30
Code: https://github.com/NUS-HPC-AI-Lab/OrScale
License: http://creativecommons.org/licenses/by/4.0/
The gist: Muon fixes the direction of every matrix-valued update at the polar factor of its momentum, while each layer's step magnitude is addressed only by a static shape correction.
Terminology
Abstract
Muon fixes the direction of every matrix-valued update at the polar factor of its momentum, while each layer's step magnitude is addressed only by a static shape correction. We derive a dynamic per-layer scalar by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information. The resulting method, OrScale, uses the norm of the parameter-space direction actually applied and anchors each layer's ratio at one via a per-layer calibration, so that the Moonlight recipe (tuned for AdamW, shared with Muon via RMS matching) transfers with no additional sweep; a component ablation confirms each design choice is individually load-bearing. Theoretically, OrScale retains a nuclear-norm O(1/sqrt T) convergence rate for any clipped multiplier and achieves a strict layer-adaptive descent gain κ eff>1 under two conditions estimable from standard training diagnostics---a bound that predicts the gain should grow with architectural heterogeneity. Experiments confirm the prediction: with every hyperparameter inherited verbatim from the Moonlight recipe, OrScale matches or beats Muon+Moonlight across dense 125M--1.1B FineWeb-Edu pre-training, and on a 16B-A3B mixture-of-experts model---where the logged trust ratios separate cleanly by layer class---the gap widens by an order of magnitude to 0.130 nats (3.8% relative) at parity wall-clock cost.
Sources
- Modular Duality in Deep Learning
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Adam: A Method for Stochastic Optimization
- One weird trick for parallelizing convolutional neural networks
- Muon is Scalable for LLM Training
- Decoupled Weight Decay Regularization
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- AdaMuon: Adaptive Muon Optimizer
- Kimi K2.5: Visual Agentic Intelligence
- Kimi K2: Open Agentic Intelligence
- SOAP: Improving and Stabilizing Shampoo using Adam
- Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks
- Large Batch Training of Convolutional Networks
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks