ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
cs.LG, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: 23 pages, 15 figures
Code: https://github.com/Time-Rune/ExFold-MoE
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation.
Terminology
Abstract
Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts' contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.
Sources
- OpenCompass: A Universal Evaluation Platform for Large Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
- Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
- Harder Tasks Need More Experts: Dynamic Routing in MoE Models
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- REAM: Merging Improves Pruning of Experts in LLMs
- Mixtral of Experts
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
- Generalizing Verifiable Instruction Following
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks