Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion large language models (dLLMs) offer a promising alternative to autoregressive generation by decoding multiple tokens in parallel through iterative denoising.
Terminology
Abstract
Diffusion large language models (dLLMs) offer a promising alternative to autoregressive generation by decoding multiple tokens in parallel through iterative denoising. However, increasing decoding parallelism often degrades generation quality, as early errors can contaminate later contexts. Revocable decoding mitigates this issue by re-evaluating decoded tokens and remasking unreliable ones, but existing methods overlook that unreliable tokens may also corrupt the verification context itself. We identify this failure mode and propose Dependency-Aware Revocable Decoding (DARD), a training-free framework that separates tokens into masked, candidate, and unmasked states. DARD verifies candidate tokens using a selective context that excludes less reliable tokens and adaptively regulates their influence on subsequent decoding. Experiments across 12 textual and multimodal benchmarks on 3 open-source dLLMs show that DARD consistently improves the speed-quality Pareto frontier over recent revocable decoding methods, achieving a 2.71 times speedup and a 4.35-point CIDEr score gain over Saber on Flickr30K.
Sources
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- dParallel: Learnable Parallel Decoding for dLLMs
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model
- GPT-4 Technical Report
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
- Stream of Search (SoS): Learning to Search in Language
- Program Synthesis with Large Language Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Let's Verify Step by Step
- Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- DeepSeek-V3 Technical Report
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering