H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache
cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 22 pages, 12 figures
Code: https://github.com/vllm-project/speculators
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model.
Terminology
Abstract
Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model. Recent block diffusion drafters further reduce drafting latency by predicting multiple tokens in parallel. However, existing block drafters project target hidden states at every input position into a separate drafter-side KV cache, incurring per-request memory and KV-write overhead that grow with concurrency; directly reusing target KVs in place removes this cache but fails to sustain draft quality throughout the block. We propose a hybrid target-context injection method that complements direct target KV reuse with target hidden states only at the last input position, requiring no separate drafter-side KV cache. Building on this design, we propose H-Spec, a hybrid Mamba-attention parallel drafter that consumes the two target-context sources through complementary modules. Mamba modules are initialized with projected last-token target hidden states, while attention modules reuse target KVs in place. Despite its recurrent formulation, Mamba's parallel scan allows H-Spec to preserve block-parallel drafting. Across three target models and diverse tasks, H-Spec improves over the best baseline by 5.0--13.3% in mean accepted length and 5.3--12.6% in batch-size-1 inter-token latency speedup. Under concurrent serving, H-Spec consistently achieves higher throughput while maintaining lower KV cache utilization than baselines across evaluated concurrency levels.
Sources
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Recurrent Drafter for Fast Speculative Decoding in Large Language Models
- VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
- The Llama 3 Herd of Models
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
- DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
- D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
- Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks