Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation
Gongli Zhang, Zhulin Liu, C. L. Philip Chen
South China University of Technology · Pazhou Lab · Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human
cs.LG, cs.CL, cs.CV
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 11 pages, 6 figures, and 7 tables; includes supplementary material. Code is available at https://github.com/existence0420/UniF-MoE
Code: https://github.com/existence0420/UniF-MoE
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts.
Terminology
Summary
Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers adapt the number of active experts. However, these decisions are usually made independently, overlooking a basic dependency: extracting reusable computation changes both what remains and how much expert capacity the remainder needs.
The paper studies this dependency by decomposing sparsely upcycled feed-forward experts into key-value channels. The authors find that co-activated experts align at a subset of value positions; removing these positions changes expert preference; and greater shared coverage is associated with lower residual expert demand.
The paper presents three main observations from diagnostic experiments on a standard top-2 MoE model trained on TerraIncognita:
Observation 1: Co-activation identifies reusable responses. The Pearson correlation between value alignment (ρij) and co-activation frequency is 0.697. Frequently co-activated pairs align at up to roughly 80% of their value positions, while some rarely selected pairs align at less than 2%.
Observation 2: Residual computation needs a new routing decision. After extracting aligned positions, only 5.7% of tokens keep the same top-2 set, and the mean diagonal mass is 26.9%. The original router scores the complete response, not the residual alone.
Observation 3: Residual demand varies across tokens. Shared ratio and mean residual demand correlate at −0.673 (bootstrap 95% CI [−0.899, −0.315]). At m=2, pair-balanced success rates are 56.3% for high-shared pairs and 11.8% for low-shared pairs.
These observations lead to one principle: share first, then route what remains.
The paper argues that shared modeling, fine-grained computation, and dynamic routing are three stages of one allocation problem, not three parallel knobs.
The paper instantiates this principle in UniF-MoE, a unified framework for token-adaptive MoE computation with the following components:
Each layer contains one shared expert Eshr and K residual experts, all divided into B aligned blocks of width M = H/B. No hidden position is assigned to both pathways for the same token.
A shared-demand score α(x) = τ + (1 − 2τ)σ(xWshr) with τ = (B−1)/B2 controls both the number of shared blocks b(x) = round(Bα(x)) and the mixture weight. Every token uses at least one shared block and leaves at least one block for residual computation.
Key prototypes µb (mean of up-projection keys) select which blocks provide reusable content via priority scores ub(x) = xµb.
The residual demand is β(x) = 1 − α(x). The number of active residual experts is determined by cumulative routing mass: k(x) = min n ∈ 1,...,K: Σi=1n pi(x) ≥ β(x). Equation (27) makes expert count conditional on the preceding shared allocation.
The final output is y = α(x)Eˢʰʳ(x) + Σi=1k(ˣ) pi(x)EiR(x), with controlled overshoot: 1 ≤ α(x) + Pk(x)(x) < 1 + pk(x)(x).
A Gram regularizer Ldiv = (Wg⋆)TWg⋆ − I K+1 F separates and normalizes router embeddings. The total loss is L = Ltask + λdiv·Ldiv.
UniF-MoE achieves the best average result (69.5%), leading on PACS (89.6), VLCS (81.7), and DomainNet (49.4), and tying the best result on OfficeHome (74.2). It outperforms dense DeiT-S/16, static MoEs (GMoE, EMoE, EMoE-L), domain-generalization methods (LFME, DMDA, PC-MoE), and dynamic MoEs (DynMoE, MASS).
UniF-MoE performs best on all five GLUE tasks with average 82.76, surpassing the per-task oracle envelope of fixed top-k variants (81.95), DynMoE (81.64), and MASS (82.19).
Relative to top-2 GMoE, UniF-MoE activates 9.1% fewer parameters and uses 16.1% fewer FLOPs, reducing inference time and memory by 45.2% and 52.7%.
It also undercuts DynMoE on every compute and runtime statistic.
-
Three adaptive decisions: Every replacement lowers accuracy and increases computation.
Fixing α is most disruptive because an incorrect shared-residual split propagates to both later decisions.
-
Diversity regularization: λdiv = 0.01 is optimal. It
reduces mean pair co-activation by 62.9% and moves the router embeddings close to orthonormality.
-
Block granularity: B = 8 or 16 performs best on four of five DomainBed datasets.
"Shared modeling, fine-grained computation, and dynamic routing need not be separate mechanisms. By exposing their shared-residual dependency, this work turns them into one ordered rule: identify reusable computation first, then route what remains. The framework
improves predictive performance while reducing activated computation and measured inference overhead."
Improvements for AI systems
Improvements to AI systems:
-
Implement a two-stage allocation pipeline in MoE architectures: Replace independent routing decisions with a sequential process where the model first identifies and extracts reusable/shared computation blocks, then routes only the remaining residual demand to experts. This directly reduces parameter activation and FLOPs while improving accuracy.
-
Add token-adaptive shared block selection: Use learned key prototypes to dynamically select which shared blocks (out of B blocks) each token uses, rather than using a fixed shared expert. This allows the model to adaptively reuse different types of knowledge per token, improving efficiency without sacrificing expressiveness.
-
Make expert count a function of residual demand: Replace fixed top-k routing with cumulative routing mass thresholding, where the number of active experts k(x) is determined by how much residual computation remains after shared extraction. This eliminates wasted computation on tokens that need fewer experts.
-
Integrate Gram-based diversity regularization on router embeddings: Add a loss term that pushes router embeddings toward orthonormality (Ldiv = (Wg*)TWg* − I F). This reduces redundant co-activation between experts by 62.9%, improving specialization and overall performance.
-
Co-design shared-residual split with routing decisions: Instead of tuning shared ratio, fine-grained computation, and dynamic routing separately, treat them as one joint optimization problem. The paper shows fixing any one of these three decisions degrades accuracy and increases compute—so they must be optimized together.
-
Use blockwise partitioning with mutual exclusion: Ensure no hidden position is assigned to both shared and residual pathways for the same token. This prevents redundant computation and forces clean separation of reusable vs. token-specific knowledge.
What the improved AI system can do:
-
Achieve higher accuracy with less compute: On GLUE, the improved system reaches 82.76 average score while activating 9.1% fewer parameters and using 16.1% fewer FLOPs than a standard top-2 MoE. It also reduces inference time by 45.2% and memory by 52.7%.
-
Adapt computation per token dynamically: The system automatically uses fewer experts for tokens with high shared coverage and more experts for tokens requiring novel computation, without manual tuning.
-
Generalize better across domains: On DomainBed benchmarks, it outperforms both static MoEs and prior dynamic routing methods (e.g., 69.5% average vs. 68.9% for the best prior method), showing improved robustness to distribution shift.
-
Avoid the
oracle envelope
problem: Unlike fixed top-k variants that require knowing the optimal k per task, the improved system automatically finds the right expert count per token, surpassing even the per-task oracle of fixed variants (82.76 vs. 81.95 on GLUE). -
Scale efficiently: The blockwise shared-residual design (B=8 or 16) provides a natural knob for controlling granularity, allowing the system to trade off between parameter efficiency and capacity as needed for different model sizes or hardware constraints.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks