OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading
cs.DC, cs.CL
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/flashserve/OLED-MoE
Terminology
Sources
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
- LLaDA2.1: Speeding Up Text Diffusion via Token Editing
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
- DFlash: Block Diffusion for Flash Speculative Decoding
- Evaluating Large Language Models Trained on Code
- KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
- Training Verifiers to Solve Math Word Problems
- Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
- ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- CDLM: Consistency Diffusion Language Models For Faster Sampling
- OverFill: Two-Stage Models for Efficient Language Model Decoding
- A Survey on Diffusion Language Models
- In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
- DeepSeek-V3 Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing