Draft-OPD: On-Policy Distillation for Speculative Draft Models
cs.CL
Submitted: 2026-05-28
Updated: 2026-09-18
Code: https://github.com/sahil280114/codealpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verified in parallel.
Terminology
Abstract
Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verified in parallel. A common way to build draft models, like EAGLE3 or DFlash is supervised fine-tuning (SFT) on target-generated trajectories. However, we observe that SFT quickly plateaus: the draft model's acceptance length on test data stops improving. The reason is an offline-to-inference mismatch: In SFT, the drafter learns from fixed target-generated trajectories, whereas during speculative decoding it is evaluated on blocks proposed under its own policy. This motivates on-policy distillation (OPD), where the target model supervises the drafter on draft-induced states. Yet OPD remains difficult for draft models, as they cannot reliably roll out complete sequences independently, whereas target-assisted generation makes the collected sequences follow the target distribution and thus eliminates the on-policy signal. We therefore propose Draft-OPD, which uses target-assisted rollout for stable continuations and replays drafting from the verification-exposed error positions. This allows the drafter to learn from target feedback on both accepted and rejected proposals, focusing training on the draft-induced errors that limit speculative acceptance. Experiments show that Draft-OPD achieves over 5 times lossless acceleration for thinking models across diverse tasks, improving over EAGLE-3 and DFlash by 23% and 13%.
Sources
- Program Synthesis with Large Language Models
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- Let's Verify Step by Step
- DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
- Online Speculative Decoding
- SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
- HybridFlow: A Flexible and Efficient RLHF Framework
- Qwen3 Technical Report
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- SGLang: Efficient Execution of Structured Language Model Programs
- DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering