DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
Yujie Zhang, Shivam Aggarwal, Tulika Mitra
cs.DC, cs.LG
Submitted: 2026-08-21
Updated: 2026-08-24
Code: https://github.com/ecolab-nus/DAOP
Terminology
Sources
- Mixtral of Experts
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Fast Inference of Mixture-of-Experts Language Models with Offloading
- MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
- EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices
- Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
- A Review of Sparse Expert Models in Deep Learning
- A Survey on Mixture of Experts in Large Language Models
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Measuring Massive Multitask Language Understanding
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing