OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

arXiv:2609.33385 · cs.DC, cs.CL · Submitted 2026-09-27 · Read on arXiv

cs.DC, cs.CL

Submitted: 2026-09-27

Updated: 2026-09-27

Code: https://github.com/flashserve/OLED-MoE

Terminology

Sources

Related papers