Federated Compositional Muon Optimizer for Matrix-Wise Models
Wang Yan, Feihu Huang
Nanjing University of Aeronautics and Astronautics
cs.LG, math.OC
Submitted: 2026-08-18
Updated: 2026-08-19
Comments: 45 pages
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 95/100
Terminology
Summary
Affiliation: College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China
arXiv: 2608.12710v1 [cs.LG] 13 Aug 2026
The paper states: "Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. To fill this gap, we propose an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems."
The authors propose two algorithms:
-
FedCoMuon - builds on
compositional gradient tracking and orthogonalized momentum
-
FedCoMuon-VR - a variance-reduced variant based on
a momentum-based variance reduced technique
The paper proves convergence properties under the non-i.i.d. and non-convex settings,
with FedCoMuon-VR achieving a lower sample complexity of O(ϵ−3) for finding an ϵ-stationary solution than the existing FedMuon algorithms.
The paper studies the following distributed matrix-wise compositional optimization problem:
W in R m times n 1 over K sum k=1 K E zeta about D f k [f k (E xi about D g k [g k(W; xi)]); zeta]
where the inner and outer expected mappings are defined as g k(W) ≜ E ξ∼D g k[g k(W; ξ)]: R m×n → R d and f k(y) ≜ E ζ∼D f k[f k(y; ζ)]: R d → R, respectively.
The global objective is F(W) = K−1 Σ k=1 K F k(W), where F k(W) = f k(g k(W)).
-
Algorithm Development: The authors
develop a class of effective Muon-type federated compositional algorithms (i.e., FedCoMuon and FedCoMuon-VR) for distributed matrix-wise stochastic compositional optimization.
FedCoMuoncombines compositional gradient tracking with orthogonalized matrix momentum, while FedCoMuon-VR further incorporates a momentum-based variance-reduction technique.
-
Convergence Analysis: The paper provides
a solid convergence analysis for the proposed algorithms under non-convex and non-i.i.d. settings.
Specifically:
-
FedCoMuon requires
a sample complexity of O(ϵ−4) and a communication complexity of O(ϵ−3)
-
FedCoMuon-VR has
a lower sample complexity of O(ϵ−3) while retaining the same communication complexity
-
FedCoMuon-VR has
a lower sample complexity than the existing Federated Muon algorithms
- Experimental Validation:
Experiments on the task-distributed meta learning problem and robust federated learning demonstrate effectiveness of the proposed algorithms.
The algorithm operates as follows:
-
Each client maintains a matrix momentum M t k updated as: M t+1 k = βZ t+1 k + (1−β)M t k
-
The stochastic compositional gradient is: Z t+1 k = ∇g k(W t+1 k; ξ t+1 k)(∇ y f k(u t+1 k; ζ t+1 k) ⊗ I n)
-
The inner function value is tracked via: u t+1 k = αg k(W t+1 k; ξ t+1 k) + (1−α)u t k
-
Orthogonalization uses
Newton–Schulz iterations instead of an expensive exact SVD
-
Every τ local iterations,
the server receives W t+1 k, M t+1 k k=1 K, averages them to obtain W̄ t+1 and M̄ t+1, and sends the averaged variables back to all clients
This variant incorporates momentum-based variance reduction:
-
Tracks inner function value u t+1 k, outer gradient v t+1 k, and inner Jacobian H t+1 k recursively
-
Uses
Euclidean projections onto B f:= v: ∥v∥ ≤ C f and B g:= H: ∥H∥ F ≤ C g
-
Updates momentum as: M t+1 k = (1−ρ)M t k + ρH t+1 k(v t+1 k ⊗ I n)
-
Each recursive estimator evaluates its function on the same fresh sample at two consecutive iterates, which reduces the stochastic estimation error
Under Assumptions 1–5, with parameters satisfying 0 < α, β < 1 and ατ ≤ 1, the convergence bound is:
1 over T sum t=0 T-1 E grad F(t) F at most F(0) - F* over eta T + 2 sqrt n C g L f sigma g over beta T + 2 sqrt n (1 over beta T + 2 beta tau(C g sigma f + C f sigma grad g)) +
With parameter choices η = T(−3/4), α = β = T(−1/2), τ = T(1/4), the bound becomes O(T(−1/4)), requiring T = O(ϵ−4) iterations, with sample complexity O(ϵ−4) and communication complexity O(ϵ−3).
Under Assumptions 1–3, 6, and 7, with η > 0, 0 0, the convergence bound is:
1 over T sum t=0 T-1 E grad F(t) F at most F(0) - F* over eta T + L F sqrt n eta over 2 + 2 sqrt n (4C g C f over rho T + L F sqrt n eta over rho) + (2 + 6 sqrt rho tau) [] 1/2 +
With parameter choices η = α = β = γ = T(−2/3), ρ = T(−1/3), b = T(2/3), τ = O(1), the bound becomes O(T(−1/3)), requiring T = O(ϵ−3) iterations, with per-client sample complexity O(ϵ−3) and communication complexity O(ϵ−3).
-
Assumption 1: The global objective F(W) has a lower bound F* > −∞
-
Assumption 2: Unbiased stochastic oracles with bounded variances (σ g, σ∇g, σ f)
-
Assumption 3: Bounded gradient moments (C g, C f)
-
Assumption 4: Smoothness of population mappings (L g, L f-smooth)
-
Assumption 5: Gradient heterogeneity bounded by δ
-
Assumption 6: Mean-square sample smoothness
-
Assumption 7: Client heterogeneity bounded by Δ g, Δ∇g, Δ f
MNIST Image Classification: FedCoMuon and FedCoMuon-VR outperform the baseline methods in terms of both training and test performance under the highly imbalanced data partition.
FedCoMuon-VR exhibits more stable convergence and achieves the best overall performance.
WikiText-2 Language Modeling: Both FedCoMuon and FedCoMuon-VR converge faster and achieve lower test loss and perplexity than the baseline methods.
FedCoMuon-VR achieves the best overall performance.
CNN-Based Meta Learning (CIFAR-10): FedCoMuon and FedCoMuon-VR achieve better overall performance than the baseline methods under different levels of data heterogeneity.
The advantage becomes more pronounced
as heterogeneity increases.
ViT-Tiny-Based Meta Learning: "FedCoMuon and FedCoMuon-VR achieve substantially higher test accuracy and lower test loss than FedMAML and the federated compositional baselines. They also remain competitive with the three FedMuon baselines throughout training."
The paper compares against:
-
Standard federated baselines: FedAvg, FedMAML
-
Muon-based federated methods: FedMuon [Zhang and Gao, 2025], FedMuon-LGA [Liu et al., 2025b], FedMuon-BC [Takezawa et al., 2026]
-
Federated compositional methods: ComFedL [Huang and Li, 2021], Local-SCGDM [Gao et al., 2022]
The paper states: "In this paper, we studied the matrix-wise composition optimization, and proposed a class of effective federated compositional Muon algorithms (i.e., FedCoMuon and FedCoMuon-VR), which build on compositional gradient tracking and orthogonalized momentum. In theory, we established convergence guarantees under non-i.i.d. and non-convex settings. In particular, our FedCoMuon-VR algorithm achieves a lower sample complexity of O(ϵ−3) for finding an ϵ-stationary solution than the existing FedMuon algorithms. Extensive experiments on robust federated learning and the task-distributed meta learning problem demonstrate the effectiveness of our proposed algorithms."
Improvements for AI systems
Based on this paper, I can make the following specific improvements to AI systems:
-
What I can do: Implement FedCoMuon's orthogonalized momentum (via Newton-Schulz iterations) to replace standard Adam/SGD optimizers when training matrix-valued parameters (e.g., weight matrices in transformers, CNNs, or embeddings) across distributed clients.
-
Improved capability: Faster convergence on matrix-wise models (e.g., ViT, GPT-style architectures) in federated settings, especially when parameters are naturally structured as matrices rather than flattened vectors.
-
What I can do: Use the compositional gradient tracking mechanism (tracking inner function values
u t kand outer gradients jointly) to optimize objectives with nested expectations, such as meta-learning (inner loop task adaptation + outer loop meta-objective) or robust learning with data-dependent regularizers. -
Improved capability: Directly optimize hierarchical problems (e.g.,
f(E[g(W)])) without resorting to biased approximations or two-level heuristics, yielding more stable training in task-distributed meta-learning and adversarial/robust federated learning. -
What I can do: Integrate FedCoMuon-VR's momentum-based variance reduction (recursive estimators for inner Jacobian
H t k, outer gradientv t k, and inner valueu t kwith Euclidean projections) to reduce stochastic noise under heterogeneous client data distributions. -
Improved capability: Achieve
O(ε−3)sample complexity (vs.O(ε−4)for standard FedMuon) in non-convex, non-IID federated settings, meaning fewer local samples needed per client to reach a target accuracy—critical for communication- and computation-limited edge devices. -
What I can do: Adopt the periodic server-client synchronization (every
τlocal steps) with averaged momentum matricesM̄ t, reducing communication rounds toO(ε−3)while maintaining convergence. -
Improved capability: Train large matrix-wise models (e.g., language models on WikiText-2) with fewer server-client exchanges, making federated training feasible over low-bandwidth networks without sacrificing model quality.
-
What I can do: Leverage the bounded gradient heterogeneity assumptions (Assumptions 5 and 7) and the algorithm's explicit handling of non-IID distributions to improve performance under imbalanced or skewed data partitions (as shown in MNIST experiments).
-
Improved capability: Maintain high test accuracy and stable convergence even when clients have vastly different data distributions (e.g., medical imaging from different hospitals), outperforming FedAvg and FedMAML baselines.
-
What I can do: Replace expensive SVD with Newton-Schulz iterations for momentum orthogonalization, reducing per-iteration computational cost while preserving the beneficial rotation-invariant properties of Muon optimizers.
-
Improved capability: Train very large matrix dimensions (e.g.,
m×nwithm,n > 104) in federated settings where exact SVD would be prohibitively expensive, enabling use on high-dimensional models like ViT-Tiny.
-
Federated Meta-Learning: Train a global model that quickly adapts to new tasks (e.g., few-shot classification on CIFAR-10) across distributed clients with heterogeneous task distributions, achieving higher accuracy than FedMAML and compositional baselines.
-
Robust Federated Learning: Train models that are resilient to label noise or adversarial perturbations in a distributed setting, with lower test loss and perplexity on language modeling tasks (WikiText-2) compared to existing federated optimizers.
-
Resource-Constrained Deployment: Operate on devices with limited local data and bandwidth, requiring fewer samples (
O(ε−3)) and communication rounds (O(ε−3)) to reach a target error, making it suitable for mobile or IoT federated learning. -
Structured Model Training: Efficiently optimize models where parameters are naturally matrix-valued (e.g., attention layers, embedding matrices) without flattening, preserving structural inductive biases and improving convergence speed.
-
Non-Convex, Non-IID Guarantees: Provide theoretical convergence guarantees in challenging settings where standard federated optimizers (e.g., FedAvg) may diverge or plateau, ensuring reliable training across heterogeneous clients.
Abstract
Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. To fill this gap, we propose an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems. Specifically, our FedCoMuon optimizer builds on compositional gradient tracking and orthogonalized momentum. Moreover, we propose a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique. In theory, we analyze the convergence properties of our algorithms under the non-i.i.d. and non-convex settings. In particular, we prove that our FedCoMuon-VR obtains a lower sample complexity of O(epsilon-3) for finding an epsilon-stationary solution than the existing FedMuon algorithms. Extensive numerical experiments on robust federated learning and task-distributed risk-sensitive meta learning show that our proposed methods are competitive with existing compositional baselines and achieve the best reported accuracy in several settings.
Sources
- On the Convergence of Muon and Beyond
- Personalized Federated Learning: A Meta-Learning Approach
- Compositional federated learning: Applications in distributionally robust averaging and meta learning
- LiMuon: Light and Fast Muon Optimizer for Large Models
- Advances and Open Problems in Federated Learning
- Convergence of Muon with Newton-Schulz
- Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization
- A Note on the Convergence of Muon
- Muon is Scalable for LLM Training
- FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
- Pointer Sentinel Mixture Models
- Muon is Provably Faster with Momentum Variance Reduction
- Communication-Efficient Gluon in Federated Learning
- Adaptive Federated Optimization
- Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- On the Convergence Analysis of Muon
- On Provable Benefits of Muon in Federated Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks