Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
cs.LG, cs.CL, cs.IT, math.IT
Submitted: 2026-08-27
Updated: 2026-09-27
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Block drafters propose several tokens in one forward pass, before earlier target tokens are realised.
Terminology
Abstract
Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches 0.286 at the final slot on Qwen3-4B, limiting even the best proposal to 71% per-slot acceptance. Second, one realised token removes 86 -- 100% of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for 43 -- 64% of DFlash rejection and 85 -- 92% of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.
Sources
- Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
- Program Synthesis with Large Language Models
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- Speculative Decoding with Big Little Decoder
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding
- Online Speculative Decoding
- The Parallelism Tradeoff: Limitations of Log-Precision Transformers
- Formal Language Recognition by Hard Attention Transformers: Perspectives from Circuit Complexity
- Gemma 4 Technical Report
- xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
- D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
- Qwen3 Technical Report
- DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks