Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
- dKV-Cache: The Cache for Diffusion Language Models
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
- dInfer: An Efficient Inference Framework for Diffusion Language Models
- TeDiServe: High SLO Attainment Serving for Diffusion Language Models
- Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
- Taming the Memory Footprint Crisis: System Design for Production Diffusion LLM Serving
- Training Verifiers to Solve Math Word Problems
- Queueing Analysis of GPU-Based Inference Servers with Dynamic Batching: A Closed-Form Characterization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection