Towards a Statistical Understanding of Mixture-of-Experts
stat.ML, cs.LG, math.ST, stat.TH
Submitted: 2026-09-03
Updated: 2026-09-03
Terminology
Sources
- DeepSeek-V3 Technical Report
- Combining pre-trained models via localized model averaging
- Mixture of A Million Experts
- Mixtral of Experts
- Kimi K3: Open Frontier Intelligence
- On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
- GPT-4 Technical Report
- Route Experts by Sequence, not by Token
- When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
- Toward Inference-optimal Mixture-of-Expert Large Language Models
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey