Less is MoE: Trimming Experts in Domain-Specialist Language Models
cs.LG, cs.CL
Submitted: 2026-06-04
Updated: 2026-09-08
Comments: To appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026), Main Conference
Journal ref: Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing, 2026
Code: https://github.com/huggingface/accelerate
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment challenges.
Terminology
Abstract
Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment challenges. Prior MoE compression approaches catastrophically fail when evaluated on general-purpose benchmarks beyond commonsense reasoning. We trace this failure to the granularity of compression: important capabilities are distributed across experts but concentrated in FFN sparse intermediate dimensions. To identify these dimensions, we use Fisher importance which outperforms activation-, router-score-, and magnitude-based alternatives, and identifies tiny sets of task-critical dimensions: in Qwen1.5-MoE, removing as few as 12 of 1.35M routed-FFN intermediate dimensions collapses GSM8K accuracy while largely preserving factual-knowledge performance. Building on this, we propose Fisher-MoE, which operates within FFN to remove intermediate dimensions ranked by Fisher importance. At the same 50% MoE compression ratio, Fisher-MoE preserves model capability, while reducing weight memory by 45% and improving inference throughput by 21%. These findings suggest intermediate dimension granularity is an effective unit for both compression and ranking where capability concentrates in MoE models.
Sources
- Program Synthesis with Large Language Models
- Delta Decompression for MoE-based LLMs Compression
- Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
- Evaluating Large Language Models Trained on Code
- Task-Specific Expert Pruning for Sparse Mixture-of-Experts
- Training Verifiers to Solve Math Word Problems
- Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
- Sparse Matrix in Large Language Model Fine-tuning
- Measuring Mathematical Problem Solving With the MATH Dataset
- STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
- Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
- OLMoE: Open Mixture-of-Experts Language Models
- SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
- ZeRO-Offload: Democratizing Billion-Scale Model Training
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
- Qwen3 Technical Report
- Qwen2 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks