Carryover Drafting: Recycling Rejected States for Speculative Decoding
cs.LG, cs.AI
Submitted: 2026-09-13
Updated: 2026-09-13
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens.
Terminology
Abstract
Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We find that these discarded hidden states generated during target forward retain useful information about future tokens that can improve subsequent drafts. However, realizing this opportunity poses two distinct challenges. At inference, recycling overhead can increase drafting latency, diminishing the speedup gained from increased acceptance length. During training, standard parallel drafter training does not produce inference-aligned rejected states, while obtaining them through sequential rollouts would sacrifice parallelism across training positions. We introduce Carryover Drafting, which addresses both challenges. Carryover recycles rejected target hidden states as temporary KV context, allowing the drafter to selectively attend to them. It reuses the drafter's existing interface and adds only a single learned embedding to distinguish rejected states from committed context. The additional KV context is replaced each drafting round, keeping its length bounded by one proposal block. We introduce parallel draft--verify--draft training that exposes the drafter to inference-aligned rejected states while preserving parallelism across training positions. Experiments with DFlash and a DSpark-derived semi-autoregressive drafter across two target models show that this simple Carryover mechanism improves average acceptance length by 6.5--14.7% and end-to-end vLLM speedup by 7.9--14.4% over the corresponding baselines, with speedup gains reaching 28.8% on translation.
Sources
- Accelerating Large Language Model Decoding with Speculative Sampling
- Make Every Draft Count: Hidden State based Speculative Decoding
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Gemma 4 Technical Report
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- Draft-OPD: On-Policy Distillation for Speculative Draft Models
- Accelerating Speculative Decoding with Block Diffusion Draft Trees
- xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
- D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
- Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks