Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis
cs.AI
Submitted: 2026-08-21
Updated: 2026-08-29
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel.
Terminology
Abstract
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.
Sources
- Pixtral 12B
- Qwen3-VL Technical Report
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
- MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
- Better & Faster Large Language Models via Multi-token Prediction
- Speculative Decoding for Autoregressive Video Generation
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
- SpecVLM: Fast Speculative Decoding in Vision-Language Models
- Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding
- LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
- LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
- Speculative Decoding Reimagined for Multimodal Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection