Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Code available at: https://github.com/VILA-Lab/Flash-dLLM
Code: https://github.com/VILA-Lab/Flash-dLLM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation.
Terminology
Abstract
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce Flash-dLLM, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger batch size. Extensive experiments on mathematical reasoning and code-generation benchmarks demonstrate that Flash-dLLM consistently outperforms existing state-of-the-art dLLM acceleration methods in both inference speed and memory efficiency. In particular, it achieves 5.1 times and 11.0 times speedups over prior strongest baseline Elastic-Cache on GSM8K and HumanEval, respectively.
Sources
- GPT-4 Technical Report
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
- Program Synthesis with Large Language Models
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- Accelerating Diffusion LLM Inference via Local Determinism Propagation
- Diffusion Language Models Know the Answer Before Decoding
- A Survey on Diffusion Language Models
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- dKV-Cache: The Cache for Diffusion Language Models
- dInfer: An Efficient Inference Framework for Diffusion Language Models
- Attention Is All You Need for KV Cache in Diffusion LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering