Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
cs.CL
Submitted: 2026-05-01
Updated: 2026-09-05
License: http://creativecommons.org/licenses/by/4.0/
The gist: Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for sparse Mixture-of-Experts (MoE) models.
Terminology
Abstract
Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for sparse Mixture-of-Experts (MoE) models. As the draft tree grows, different branches activate different experts, expanding the union of activated experts and substantially increasing target-side verification cost. We propose EVICT, a training-free, hyperparameter-free, and lossless adaptive verification method for MoE speculative decoding. EVICT makes every verified token count by truncating the draft tree before target verification and retaining only the cost-effective prefix. It leverages fine-grained drafter signals to estimate candidate benefit, combines them with offline-profiled verification cost, and remains highly compatible with the high-performance graph-based serving framework SGLang. Extensive experiments on diverse MoE backbones and benchmarks show that EVICT achieves up to 2.35x speedup over autoregressive decoding and an average 1.21x speedup over the state-of-the-art baseline EAGLE-3, while significantly reducing unnecessary expert activations during verification.
Sources
- Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Qwen3-Coder-Next Technical Report
- Accelerating Large Language Model Decoding with Speculative Sampling
- Evaluating Large Language Models Trained on Code
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Training Verifiers to Solve Math Word Problems
- Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- DeepSeek-V3 Technical Report
- TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees
- UCCL-EP: Portable Expert-Parallel Communication
- MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
- Qwen3 Technical Report
- Utility-Driven Speculative Decoding for Mixture-of-Experts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering