ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
cs.LG, cs.CL, cs.DC
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: Accepted to COLM 2026
Code: https://github.com/Amir-zsh/ASPIRE
Project page: https://amir-zsh.github.io/ASPIRE
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound.
Terminology
Abstract
Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule, even though the optimal draft length varies widely across requests and changes dynamically within each request. We propose ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components. First, a unified mixed forward allows drafting and verifying requests to coexist in the same batched forward pass, removing the need for global draft-verify phases. Second, a lightweight online speculation scheduler uses per-request acceptance-rate estimates and a batch-aware cost model to let each request independently choose when to verify. Third, an intra-draft refresh layer performs full attention at a single designated layer during drafting, updating the sparse context at every draft step to reduce staleness during drafting. Across three models and five reasoning and long-context benchmarks, ASPIRE achieves 1.70 - 4.58 times speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately 27% over the strongest prior self-speculative baselines.
Sources
- A Comprehensive Survey on Long Context Language Modeling
- CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
- Qwen3 Technical Report
- Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks