MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

arXiv:2608.11749 · cs.LG, cs.AI, stat.ML · Submitted 2026-08-12 · Read on arXiv

Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun

Beihang University · Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing · Beijing Academy of Artificial Intelligence · Renmin University of China

cs.LG, cs.AI, stat.ML

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/KunlinLyu/MOON

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning proposes a structure-aware multi-objective optimization (MOO) method that performs gradient manipulation under spectral–nuclear

Terminology

Summary

MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning proposes a structure-aware multi-objective optimization (MOO) method that performs gradient manipulation under spectral–nuclear norm geometry and uses orthonormalized updates for matrix-valued parameters, addressing the geometric mismatch in existing Euclidean-based MOO methods for modern architectures like Transformers.

Problem and Motivation:

"Most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. The paper notes that existing MOO methods such as MGDA, PCGrad, FAMO, and other variants typically flatten model parameters into a single vector and then perform gradient manipulation in this Euclidean vector space. However, modern architectures, such as Transformers, are dominated by matrix-valued parameters. Blindly flattening all parameters into vectors ignores the linear-mapping structure of matrix-valued parameters and may fail to identify the steepest descent direction under matrix geometry."

Proposed Method:

MOON performs gradient manipulation under spectral–nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. The method is derived from simultaneous matrix-smooth upper bounds of the objectives and formulates the selection of a common descent direction as a spectral-norm-regularized minimax problem. Through "spectral–nuclear norm duality, the corresponding dual problem determines the task weights by minimizing the nuclear norm of the weighted aggregate gradient, while the solution to the primal problem is given by its polar factor. The update direction is given by the polar factor W t = U t V t of the momentum-smoothed aggregate gradient, where the update magnitude is controlled separately by the learning rate."

Practical Implementation:

Practical MOON incorporates three components: (1) gradient momentum to stabilize the aggregate gradient across iterations, (2) Newton–Schulz iterations to efficiently approximate its polar factor, and (3) a single-step online update to track the dual task weights. The algorithm maintains an exponential moving average of the aggregate gradient, approximates the polar factor via Newton–Schulz iterations, and updates task weights using logits with a softmax parameterization.

Theoretical Results:

For smooth non-convex objectives, the paper establishes convergence of the averaged Pareto-stationarity measure at rates of O(T-1/2) in the deterministic setting and O(T-1/4) under stochastic gradients. Specifically, Theorem 1 states: "Suppose that each objective i: R p times q to R is L-smooth with respect to the spectral norm... Choose alpha = sqrt B over LT, beta = 1 over mHT and 0 < gamma < 1 over mHT. Then, for every fixed mu in (0,1], the iterates generated by Algorithm 1 satisfy 1 over T sum t=1 T z* t in m grad L(t)z* t S 1 at most O(T-1/2). Theorem 2 states: Under the same assumption as Theorem 1... Choosing alpha = sqrt mu B over TL, beta = 1 over(H+ sqrt m rho sigma)mT, 0 < gamma < 1 over(H+ sqrt m rho sigma)mT and mu = sqrt LB over sigma rho sqrt T, 1, where rho:= sqrt p,q, then the iterations of Algorithm 1 satisfy E [1 over T sum t=1 T z* t in m grad L(t)z* t S 1] at most O(1/T 1/4)."

Key Theoretical Insight:

The paper shows that at the same iterate, the exact MOON weighting provides a no-looser nuclear-norm-dependent certificate than Euclidean min-norm weighting, meaning G t(z t MOON) S 1 at most G t(z t E) S 1 for any Euclidean min-norm weighting z t E.

Experimental Results:

MOON is evaluated on five benchmarks: MultiMNIST (2 tasks), NYU-v2 (3 tasks), CityScapes (2 tasks), QM9 (11 tasks), and CelebA (40 tasks), comparing against 12 MOO baselines. Key results include:

  • On MultiMNIST with ViT backbone, MOON achieves the highest average accuracy of 95.65%, outperforming FAMO (95.39%) and MGDA (94.84%).

  • On NYU-v2, MOON achieves the best average performance drop m% = -4.63, outperforming all baselines.

  • On CityScapes, MOON achieves m% = 1.54, the best among compared methods.

  • On QM9, MOON achieves the lowest performance drop m% = 3.35 and lowest MAE on most individual tasks.

  • On CelebA (40 tasks), MOON achieves m% = 4.65, outperforming all baselines.

Additional Findings:

  • A toy experiment shows MOON converges faster and reaches lower final average loss than MGDA and MGDA+Muon.

  • Ablation studies show that both orthonormalized updates and momentum are important components.

  • MOON improves wall-clock convergence (reaching a target loss 39.8% faster than MGDA and 14.1% faster than FAMO) with negligible memory overhead.

  • MOON also demonstrates effectiveness in multi-objective reinforcement fine-tuning of a large language model (Qwen3-1.7B-Base).

Conclusion:

"This work addresses the geometric mismatch between conventional Euclidean MOO methods and the matrix-valued parameters prevalent in modern architectures such as Transformers. We proposed MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral–nuclear norm geometry and constructs matrix-aware orthonormalized updates... Experiments demonstrate that MOON improves optimization efficiency while achieving competitive or improved final performance."

Improvements for AI systems

Improvements to AI Systems:

  1. Matrix-Geometry-Aware Multi-Task Optimization: Replace Euclidean-based gradient manipulation (used in MGDA, PCGrad, FAMO) with spectral–nuclear norm geometry for all matrix-valued parameters (e.g., attention weights, feed-forward layers in Transformers). This yields steeper descent directions, accelerating convergence by 39.8% over MGDA and 14.1% over FAMO in wall-clock time, with negligible memory overhead.

  2. Orthonormalized Update Rule for Parameter Matrices: Use the polar factor W t = U t V t of the momentum-smoothed aggregate gradient as the update direction, instead of raw Euclidean gradients. This preserves the linear-mapping structure of matrices, preventing distortion from flattening and improving optimization stability in deep networks.

  3. Adaptive Task Weighting via Nuclear-Norm Minimization: Dynamically compute task weights by minimizing the nuclear norm of the weighted aggregate gradient (dual problem), rather than using Euclidean min-norm weights. This provides a provably no-looser certificate (G t(z t MOON) S 1 at most G t(z t E) S 1), leading to more balanced multi-task performance.

  4. Stabilized Training with Gradient Momentum and Newton–Schulz Approximations: Incorporate exponential moving average of aggregate gradients to reduce variance and use efficient Newton–Schulz iterations for polar factor computation. This enables robust training under stochastic gradients, with theoretical convergence guarantees of O(T-1/4) in stochastic settings.

  5. Scalable Multi-Task Learning for High-Dimensional Outputs: Apply MOON to systems with many tasks (e.g., 40 tasks in CelebA) and large models (e.g., Qwen3-1.7B-Base for multi-objective RL fine-tuning), achieving superior average performance drops (m% = 4.65 on CelebA, best among 12 baselines) while maintaining low per-task error.

What the Improved AI System Can Do:

  • Train multi-task Transformers and LLMs with faster convergence and better final performance across diverse objectives (e.g., vision, language, multi-label classification) without task-balancing hyperparameter tuning.

  • Handle matrix-structured parameters natively (e.g., in attention heads, dense layers) without flattening, improving gradient fidelity and optimization efficiency.

  • Achieve state-of-the-art results on benchmarks like MultiMNIST (95.65% accuracy), NYU-v2 (best average performance drop-4.63), CityScapes (1.54), QM9 (3.35), and CelebA (4.65), outperforming 12 existing MOO methods.

  • Stably fine-tune large language models under multiple conflicting objectives (e.g., helpfulness, safety, style) with theoretical convergence guarantees, enabling reliable multi-objective RL alignment.

  • Reduce computational cost by using efficient polar factor approximations and momentum, making it practical for real-time or resource-constrained multi-task deployment.

Sources

Related papers