DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
Amirmohammad Karimi, Chao Gao, Negar Hassanpour
cs.CL
Submitted: 2026-08-05
Code: https://github.com/sgl-project/specforge
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them.
Terminology
Abstract
Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent block and diffusion-style drafters can predict several positions in a single pass, their training and sampling procedures are typically optimized for greedy decoding or assume that positions in the draft block are conditionally independent. This assumption becomes brittle in non-greedy speculative decoding, where the target distribution is deliberately stochastic and multiple continuations become plausible. We study this mismatch for block diffusion drafters and show that the accepted draft length degrades as the entropy of the target sampling distribution increases. We propose a dependent block drafter based on a low-rank latent mixture over token positions, complemented by an acceptance-oriented training objective that directly targets the expected verified length. Experiments with Qwen3-4B and Qwen3-8B on GSM8K, MT-Bench, HumanEval, and creative-writing benchmarks show that our approach, namely DBLast, consistently improves accepted length over independent block sampling, especially in higher-entropy decoding regimes.
Sources
- Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- Evaluating Large Language Models Trained on Code
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Training Verifiers to Solve Math Word Problems
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- Understanding R1-Zero-Like Training: A Critical Perspective
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering