OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

arXiv:2605.07815 · cs.LG, cs.CL · Submitted 2026-05-08 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-05-08

Updated: 2026-08-30

Code: https://github.com/NUS-HPC-AI-Lab/OrScale

License: http://creativecommons.org/licenses/by/4.0/

The gist: Muon fixes the direction of every matrix-valued update at the polar factor of its momentum, while each layer's step magnitude is addressed only by a static shape correction.

Terminology

Abstract

Muon fixes the direction of every matrix-valued update at the polar factor of its momentum, while each layer's step magnitude is addressed only by a static shape correction. We derive a dynamic per-layer scalar by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information. The resulting method, OrScale, uses the norm of the parameter-space direction actually applied and anchors each layer's ratio at one via a per-layer calibration, so that the Moonlight recipe (tuned for AdamW, shared with Muon via RMS matching) transfers with no additional sweep; a component ablation confirms each design choice is individually load-bearing. Theoretically, OrScale retains a nuclear-norm O(1/sqrt T) convergence rate for any clipped multiplier and achieves a strict layer-adaptive descent gain κ eff>1 under two conditions estimable from standard training diagnostics---a bound that predicts the gain should grow with architectural heterogeneity. Experiments confirm the prediction: with every hyperparameter inherited verbatim from the Moonlight recipe, OrScale matches or beats Muon+Moonlight across dense 125M--1.1B FineWeb-Edu pre-training, and on a 16B-A3B mixture-of-experts model---where the logged trust ratios separate cleanly by layer class---the gap widens by an order of magnitude to 0.130 nats (3.8% relative) at parity wall-clock cost.

Sources

Related papers