MineDraft: A Framework for Batch Parallel Speculative Decoding
cs.CL, cs.AI, cs.DC, cs.LG
Submitted: 2026-02-24
Updated: 2026-09-01
Comments: Accepted at ICML 2026
Code: https://github.com/huggingface/transformers
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model.
Terminology
Abstract
Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model. However, the performance of standard SD is often limited by the strictly sequential execution of these drafting and verification stages. To address this, this paper proposes MineDraft, a batch parallel speculative decoding (PSD) framework designed to effectively hide drafting latency by overlapping it with verification. Our theoretical analysis shows that PSD is substantially more efficient than standard SD. MineDraft realizes the PSD through a novel batch-parallel design that maintains two batches of requests, overlapping drafting for one batch with verification for the other. Our experimental results show significant improvements of in both throughput (up to 75%) and end-to-end latency (up to 39%) over standard SD. Furthermore, we have implemented MineDraft as a plugin for vLLM, demonstrating its practicality for production-ready inference systems.
Sources
- Recurrent Drafter for Fast Speculative Decoding in Large Language Models
- Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
- The Llama 3 Herd of Models
- A General Theory of Computational Scalability Based on Rational Functions
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
- TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
- ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
- Qwen3 Technical Report
- Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering