JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
cs.CL
Submitted: 2026-06-16
Updated: 2026-10-03
Code: https://github.com/hao-ai-lab/JetSpec
Terminology
Sources
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
- Program Synthesis with Large Language Models
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3 Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Better & Faster Large Language Models via Multi-token Prediction
- Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
- Accelerating Speculative Decoding with Block Diffusion Draft Trees
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering