Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
cs.CL
Submitted: 2026-07-03
Updated: 2026-09-28
Code: https://github.com/sgl-project/SpecForge
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel.
Terminology
Abstract
Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
- Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- Evaluating Large Language Models Trained on Code
- Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
- The Llama 3 Herd of Models
- AutoJudge: Judge Decoding Without Manual Annotation
- REST: Retrieval-Based Speculative Decoding
- SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
- C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
- Gemma 3 Technical Report
- OpenVLA: An Open-Source Vision-Language-Action Model
- TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees
- Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering