A Survey on Inference Optimization Techniques for Mixture of Experts Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "A Survey on Inference Optimization Techniques for Mixture of Experts Models".
Jane: The paper was written by Jiacheng Liu, Peng Tang, Wenfeng Wang, Yuhang Ren, Xiaofeng Hou et al. from Chinese University of Hong Kong and Shanghai Jiao Tong University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. I'm Tom, and alongside me is the brilliant Jane. We're diving into a paper that's been making waves in the AI world, and it's called "A Survey on Inference Optimization Techniques for Mixture of Experts Models."
Jane: And Tom, I have to say, this title is a mouthful, but it's tackling something huge. It's about making these massive AI models actually usable. You know, models like the ones that power chatbots? They're getting so big that just running them is a challenge.
Tom: Right, and the paper is all about "Mixture of Experts." Can you break that down for our listeners, Jane? What does that even mean?
Jane: So, imagine you have a huge team of specialists, but instead of everyone working on every single task, you have a smart manager who looks at each question and says, "Okay, you're a math expert, you handle this. You're a history expert, you take this one." That's the core idea. The model has all these "experts," but it only turns on a few of them for any given piece of text.
Tom: So it's like having a giant library, but you only pull out the books you need for a specific question. That's clever. But the paper's title says "inference optimization." So the problem isn't just building these models, it's actually using them in the real world?
Jane: Exactly. You can build this massive, super-smart model, but when you try to run it on a server, it's slow and it eats up a ton of memory. The paper is a survey, which means it's a map of all the different tricks and techniques researchers have come up with to make this process faster and cheaper.
Tom: A map of all the tricks. I love that. So who wrote this? Who's giving us this map?
Jane: It's a big collaboration, mostly from Shanghai Jiao Tong University and the Chinese University of Hong Kong. There are a lot of authors, but the lead researchers are Jiacheng Liu, Peng Tang, and the corresponding authors, Xiaofeng Hou and Chao Li. They've clearly been deep in this field.
Tom: And they've put together a framework to organize all these optimization techniques. It's not just a random list, right?
Jane: No, they've broken it down into three main levels. There's the model level, which is about changing the architecture of the AI itself. Then there's the system level, which is about the software and how you schedule tasks on the hardware. And finally, there's the hardware level, which is about designing new chips that are better suited for this kind of work.
Tom: Three levels. Model, system, hardware. That's a clean way to think about it. So, is the main takeaway that these MoE models are just too hard to run, or is the paper optimistic?
Jane: Oh, it's very optimistic. It shows that there's a lot of active research and many clever solutions. But it also highlights that we're not there yet. It's a field that's moving incredibly fast, and this survey is trying to keep up with it.
Tom: So it’s a snapshot of a moving target. Now, I’m curious about the actual techniques. What kind of "tricks" are we talking about? Are we shrinking the models, or are we just getting better at running them?
Jane: Both, actually. And that's what we're going to dig into next. Some people are trying to compress the experts, some are trying to offload them to different types of memory, and some are even trying to design new hardware from scratch. It's a whole ecosystem of solutions.
Tom: A whole ecosystem. I like that. So, Jane, you mentioned compressing and offloading. Let's get into the details. That's where the real meat of this survey is.
Summary: Tom: So we're back, and we're still talking about "A Survey on Inference Optimization Techniques for Mixture of Experts Models." Jane, you said the paper is a map. Let's talk about what's actually on that map.
Jane: Right. The survey starts by laying out the fundamentals. It explains that in a standard Transformer model, which is the backbone of most modern AI, there's a part called the Feed-Forward Network, or FFN. In a Mixture of Experts model, you replace that single FFN with a bunch of smaller ones—the experts.
Tom: And a router decides which of those smaller ones to use for each word, or token, in the input. That's the "sparse activation" part. It's brilliant because you get a model with a huge number of parameters, but you only compute with a small fraction of them for any given word.
Jane: Exactly. And the paper gives some great examples. They mention Mixtral eight times 7B, which has forty-six point seven billion total parameters, but only uses about thirteen billion for each token. And then there's DeepSeek-V3, which is a monster with six hundred seventy-one billion total parameters but only activates thirty-seven billion. That's the power of MoE.
Tom: So you get the brainpower of a 671B model, but the computational cost of something much smaller. That's the promise. But the survey says the reality is more complicated. What are the actual problems?
Jane: The main problem is that the "sparse" part is hard on the hardware. When you have a dense model, every parameter is used, so you can load it all into memory and just compute. With MoE, you have to constantly be moving experts in and out of memory, and you have to communicate which tokens are going to which expert across different GPUs. That communication is a huge bottleneck.
Tom: It's like having a team of specialists in different rooms. The manager has to send each task to the right room, and then the answers have to come all the way back. That travel time can eat up all the time you saved by not having everyone work on everything.
Jane: That's the perfect analogy. And the survey categorizes the solutions to this problem. At the model level, they talk about things like pruning, which is cutting out experts that aren't useful, and quantization, which is reducing the precision of the numbers to make them smaller and faster to move around.
Tom: And at the system level?
Jane: That's where they talk about expert parallelism, which is how you distribute the experts across multiple GPUs. And they talk about expert offloading, which is when you can't fit everything on the GPU, so you store some experts in the CPU's memory or even on an SSD and load them in only when they're needed.
Tom: So you're basically swapping experts in and out like a computer swaps programs in and out of RAM. That's a clever way to deal with the memory problem.
Jane: And the survey goes into a lot of detail on the different caching and prefetching strategies to make that swapping as efficient as possible. It's a really comprehensive look at the whole field.
Tom: It sounds like the paper is saying, "Here's the problem, and here are all the ways people are trying to solve it." But is there a consensus on the best approach?
Jane: No, and that's part of the point. The best approach depends on your hardware, your model, and your specific needs. Are you running on a giant server farm or on a single phone? The survey is a guide to help you pick the right tool for the job.
Tom: So it's not a one-size-fits-all solution. It's a menu of options. I'm really curious about that hardware level, though. The paper mentioned designing new chips. That's a whole different ballgame.
Improvements: Tom: We're back, and we're still on "A Survey on Inference Optimization Techniques for Mixture of Experts Models." Jane, you mentioned the paper talks about designing new hardware. That sounds like the most radical solution. Can we dig into that?
Jane: Absolutely. The survey points out that our current GPUs are really good at dense, predictable computation. But MoE is sparse and dynamic. So there's a mismatch. The paper highlights a few hardware projects that are trying to fix this. There's one called MoNDE, which uses a concept called "near-data processing."
Tom: Near-data processing. So instead of moving all the expert data to the processor, you move the processing to the data? Like, doing the math right where the memory is?
Jane: Exactly. They use a special memory type called LPDDR, which is very energy-efficient, and they put special processing cores right next to it. This way, the "cold" experts that aren't used often can be computed in memory, without having to be moved to the GPU at all. It saves a ton of bandwidth.
Tom: That's wild. So the GPU handles the popular experts, and the memory itself handles the rarely used ones. That's a real co-design between the software and the hardware. Are there other examples?
Jane: There's another one called FLAME, which is for FPGAs. FPGAs are chips you can reconfigure after they're built. FLAME uses a technique to predict which experts will be needed next, so they can be pre-loaded. It's like a hardware-level version of the prefetching we talked about at the system level.
Tom: So it's not just about the model or the software, but the entire stack has to be rethought. That's a huge undertaking. But I'm wondering, is this just about making things faster, or is there a bigger goal here?
Jane: The bigger goal is accessibility. If we can make MoE inference efficient, we can run these powerful models on smaller, cheaper hardware. That means better AI on your phone, in your car, or in a doctor's office, without needing to connect to a massive cloud server.
Tom: That would be a game-changer for privacy and for places with poor internet connections. The survey is really about democratizing access to these huge models.
Jane: And it's not just about hardware. The survey also stresses the need for better software frameworks and standardized benchmarks. Right now, it's hard to compare different optimization techniques because everyone uses different setups.
Tom: So we need a common yardstick. That makes sense. It's hard to know what the best solution is if everyone is measuring with a different ruler.
Jane: Exactly. And that's one of the key "future directions" the paper identifies. It's not just about inventing new techniques, but also about creating the infrastructure to evaluate them fairly and share them with the community.
Tom: So the paper is a call to action. It's saying, "Here's the state of the art, and here's what we need to do next." I'm curious, though, about the practical impact. Who is going to benefit from this first?
First Page: Tom: So we're back for another segment on "A Survey on Inference Optimization Techniques for Mixture of Experts Models." And I want to go back to the very first page of this paper, because it sets the stage so well.
Jane: It really does. The first page immediately connects the rise of MoE to the need for this survey. It mentions models like GPT-four and Claude, and how their massive scale is what makes them so capable. But it also points out that this scale creates huge challenges for deployment.
Tom: Right, it's that classic problem. You can build the smartest model in the world, but if it takes a minute to generate each word, or if it requires a supercomputer to run, it's not very useful in the real world. The paper calls this a "significant challenge in terms of computational efficiency and resource utilization."
Jane: And then it introduces the MoE architecture as the "promising architectural solution." The paper talks about how MoE allows you to have a model with a massive number of parameters, but only use a fraction of them for each input. That's the "conditional computation" idea.
Tom: And the first page has that great table listing all the recent MoE models. You can see the progression from NLLB and Mixtral to the massive DeepSeek-V3. It really shows how quickly this field is moving.
Jane: It does. And that table is a great example of the paper's value. It's not just theoretical. It's grounded in the real models that are being released. They even list the number of experts, the hidden dimensions, and the affiliations of the teams that built them.
Tom: So it's a snapshot of the entire MoE landscape. And the paper's introduction makes a really important point. It says that the "dynamic nature of expert activation patterns introduces complexity in resource management and scheduling that is not present in traditional dense models."
Jane: That's the core problem. In a dense model, the computation is predictable. In MoE, it depends on the input. One word might need experts one four and seven. The next word might need experts two and nine. That unpredictability is what makes it so hard to optimize.
Tom: So the paper is saying, "We have this great new architecture, but our existing tools and systems weren't built for it." And that's why we need this survey, to map out all the ways people are adapting.
Jane: And the survey is very thorough. It covers everything from tweaking the model's internal architecture to designing brand new computer chips. It's a full-stack problem, and the paper treats it that way.
Tom: I think that's what makes this paper so important. It's not just a collection of techniques. It's a framework for thinking about the problem. It helps you understand where a particular optimization fits into the bigger picture.
Jane: And that framework is what we've been talking about: model-level, system-level, and hardware-level. It gives researchers a common language to discuss their work.
Tom: A common language and a clear map. So, we've covered a lot of ground. We've talked about the problem, the solutions, and the future directions. What's the final takeaway for our listeners?
Conclusion: Tom: Well, we've reached the end of our discussion on "A Survey on Inference Optimization Techniques for Mixture of Experts Models." Jane, it's been a fantastic conversation. Can you give us a final summary?
Jane: I'd love to, Tom. This paper is essentially a field guide for anyone trying to deploy a Mixture of Experts model. It tells you that the biggest challenge isn't building the model, it's running it efficiently. And it gives you a clear, three-part framework for tackling that challenge.
Tom: The model level, the system level, and the hardware level. And at each level, there are clever tricks. You can prune and quantize the model, you can offload and cache experts at the system level, and you can even design new chips that are built for sparse computation.
Jane: Exactly. And the paper doesn't just list these techniques. It also highlights the open challenges. We need better software frameworks that natively support MoE, we need standardized benchmarks so we can compare different approaches, and we need to think about energy efficiency, not just speed.
Tom: Energy efficiency. That's a big one. It's not just about making it fast, but making it sustainable. The paper really pushes for a more holistic view of optimization.
Jane: It does. It's about making these powerful models accessible to everyone, whether that's on a huge server cluster or on a tiny edge device. And it's about doing that in a way that's responsible and sustainable.
Tom: So, for our listeners, what's the one thing you should remember about this paper?
Jane: I think it's that Mixture of Experts is a powerful idea, but it's not a free lunch. The sparsity that makes it efficient on paper creates new problems in practice. But the research community is on it, and this survey is a testament to the incredible ingenuity being applied to solve those problems.
Tom: It's a dynamic field, and this paper is a fantastic snapshot of it. We'll be watching to see what new techniques come out. Thanks for joining us, and we'll see you next time on the arXiv channel.
Jane: Goodbye, everyone!
Jiacheng Liu, Peng Tang, Wenfeng Wang, Yuhang Ren, Xiaofeng Hou, Pheng-Ann Heng, Minyi Guo, Chao Li
Chinese University of Hong Kong · Shanghai Jiao Tong University
cs.LG, cs.AI, cs.DC
Submitted: 2025-01-22
Updated: 2026-08-11
Comments: Under Review
Journal ref: ACM CSUR 2026
DOI: 10.1145/3794845
Code: https://github.com/MoE-Inf/awesome-moe-inference
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 74/100
The gist: This comprehensive survey analyzes optimization techniques for Mixture of Experts (MoE) models across the entire system stack.
Key concepts
- Mixture of Experts (MoE)
- An AI architecture where a single task is handled by multiple specialized 'experts.' Instead of using all experts for every input, a manager selects only a few relevant experts to process the text, allowing for massive models with controlled computational cost.
- Inference Optimization
- The techniques used to make large AI models run quickly and efficiently in the real world. Since MoE models are 'sparse' (only using parts of the model), optimization focuses on managing data movement and computation bottlenecks to save time and memory.
- Model Level Optimization
- Changes made directly to the AI model's architecture. Techniques include pruning (cutting out useless experts) or quantization (reducing number precision) to make the model smaller and faster while retaining performance.
- Hardware Level Optimization
- Designing specialized computer chips, such as using 'near-data processing.' This moves computation closer to the memory where data resides, reducing the massive bandwidth bottleneck caused by moving expert data.
Terminology
Summary
This comprehensive survey analyzes optimization techniques for Mixture of Experts (MoE) models across the entire system stack. The authors establish a taxonomical framework that categorizes optimization approaches into model-level, system-level, and hardware-level optimizations.
Model-Level Optimizations
Model-level optimizations aim to enhance the inherent structure and efficiency of MoE models through systematic improvements in architecture, parameter optimization, and algorithmic design.
These are categorized into three main areas:
-
Efficient Model Architecture Design: This focuses on
developing more efficient expert and attention structures.
Research includesMoE-based Attention Design,
such as MAE, MoA, and SwitchHead, andMoE-based FFN Design,
including MoE++, SCoMoE, and Pre-gated MoE. -
Model Compression Techniques: These
aim to reduce model size and memory footprint.
-
Expert Pruning: Methods are
generally divided into structured and unstructured pruning,
involving the direct removal of unimportant experts or merging groups of experts into a single expert. -
Expert Quantization: This involves
converting high-precision weights into low-precision
to reduce model size, with research focusing on quantizing expert weights (e.g., MC-MoE, QMoE, CMoE). -
Expert Distillation: A
promising solution to generate a compact, high-performance model from the original MoE model
by transferring knowledge from a large MoE teacher to a smaller student. -
Expert Decomposition: Utilizing
low-rank decomposition
to reduce model size bydecomposing a large weight matrix into smaller matrices.
-
Algorithm Improvement: This concentrates on
enhancing the dynamic aspects of MoE models.
-
Dynamic Gating: Strategies that
dynamically determine the number of experts required for each token
to exploit sparsity. -
Sparse to Dense: Approaches that
transform MoE models into dense target models all at once
to achieve maximum sparsity while maintaining performance.
System-Level Optimizations
Due to the unique structure of MoE, many studies focus on accelerating inference at the system level by leveraging the sparse activation patterns inherent to this architecture.
-
Expert Parallelism: A
key technique for deploying large MoE models across multiple devices to enable distributed execution.
Optimization occurs across four dimensions:parallelism strategy design,
load balancing
(to address disparities where some experts receive more tokens than others),all-to-all communication optimization
(to address the significant bottleneck in MoE layer execution), andtask scheduling
(tooverlap communication and computation tasks
). -
Expert Offloading: Designed for
edge devices
wherelimited GPU memory often cannot accommodate all parameters of an MoE model.
This techniqueselectively loads only the required experts
to reduce loading latency. It includes: -
Expert Prefetching:
Predicting the experts needed for subsequent layers
to minimize waiting time. -
Expert Caching: Optimizing
cache management policy... to improve the cache hit ratio and reducing transfer time.
-
Expert Loading: Aiming to
directly reduce loading costs
through methods like using low-precision experts. -
CPU Assisting: Leveraging the CPU to
assist in the process,
such as performing computation on the CPU when experts are unavailable in GPU memory.
Hardware-Level Optimizations
Recent hardware optimizations for MoE inference have addressed key challenges through novel architectures and co-design approaches.
These optimizations target critical issues such as operations per byte (Op/B) efficiency, heterogeneous computing units, and memory access patterns.
Examples include near-data processing (NDP) solutions like MoNDE, FPGA-based frameworks like FLAME, and Processing-in-Memory (PIM) architectures like Duplex.
Future Directions and Open Challenges
The survey identifies several critical challenges and opportunities:
-
Computing Infrastructure Optimization: This includes
Hardware Integration and Acceleration
(developing specialized circuits for routing and memory hierarchies for sparse access) andSystem Software Optimization
(redesigning memory management and resource allocation for dynamic expert activation). -
Operational Challenges and Requirements: This covers
Energy Efficiency and Sustainability
(considering carbon emissions and energy consumption), andLatency and Quality of Service
(addressing the difficulty of providingconsistent latency guarantees
due to variable execution paths). -
Development Support Ecosystem: This involves improving
Open-Source Frameworks
(adding native support for MoE-specific operations) andBenchmarking and Standardization
(creatingcomprehensive benchmarking and standardization frameworks
that account for the unique characteristics of MoE architectures).
Improvements for AI systems
Dynamic Token-to-Expert Routing
The system can adjust the number of active experts per token in real-time, drastically reducing computational FLOPs and energy consumption for simple tokens while allocating higher capacity to complex reasoning tasks.
Predictive Expert Prefetching and Caching
The system can anticipate which experts will be required for subsequent layers or tokens, loading them from CPU memory to GPU memory in advance, allowing massive MoE models to run on memory-constrained edge devices with near-instantaneous response times.
Asynchronous Communication-Computation Scheduling
The system can overlap all-to-all
communication tasks with active expert computation, eliminating the communication bottleneck in distributed environments and maximizing GPU utilization during large-scale inference.
Hybrid Quantized-Decomposed Expert Modules
The system can compress expert weights using a combination of low-rank decomposition and low-precision quantization, enabling the deployment of trillion-parameter models within the limited VRAM of consumer-grade hardware without significant loss in accuracy.
PIM-Integrated Sparse Routing
The system can perform expert selection and routing directly within Processing-in-Memory (PIM) architectures, bypassing the traditional memory bandwidth bottleneck and significantly increasing operations-per-byte (Op/B) efficiency.
Deterministic Latency QoS Controller
The system can implement specialized resource allocation and task scheduling that provides consistent latency guarantees, ensuring that the variable execution paths of sparse MoE models do not cause performance jitter in real-time applications like autonomous driving or voice assistants.
Energy-Aware Expert Selection
The system can utilize a routing algorithm that selects the most energy-efficient expert paths for non-critical tasks, reducing the overall carbon footprint and operational costs of large-scale AI deployments.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Efficient Large Scale Language Modeling with Mixtures of Experts
- Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
- Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
- Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
- Task-Specific Expert Pruning for Sparse Mixture-of-Experts
- A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
- Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Innovating for Tomorrow: The Convergence of SE and Green AI
- MoEUT: Mixture-of-Experts Universal Transformers
- SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek-V3 Technical Report
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Learning Factored Representations in a Deep Mixture of Experts
- Fast Inference of Mixture-of-Experts Language Models with Offloading
- LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks