Higher-order pruning of experts in mixture-of-experts language models
cs.LG, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-24
Code: https://github.com/awslabs/hybrid-model-factory
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck.
Terminology
Abstract
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.
Sources
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- DeepSeek-V3 Technical Report
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks