AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
cs.AI, cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: EMNLP Main Conference 2026
Code: https://github.com/huggingface/smolagents
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions.
Terminology
Abstract
Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this symmetry: a lightweight drafter reads the full input while the large verifier operates on the compressed view. The drafter steers the verifier via a contrastive δ-fusion of logits, modulated by a divergence-aware acceptance gate that preserves verification stability and high draft acceptance rates. Evaluated across four agentic capabilities and two end-to-end agent benchmarks, AsymSpec reaches about 90% of full-context accuracy on average, delivering 1.3 -- 1.7 times throughput speedups at 0.2 -- 0.3 times the compute cost on isolated text capabilities. These results show that asymmetric context access yields substantial gains precisely when compression discards critical reasoning signals.
Sources
- Accelerating Large Language Model Decoding with Speculative Sampling
- Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
- Retrieval-Augmented Generation for Large Language Models: A Survey
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
- Scaling Laws for Neural Language Models
- In-context Autoencoder for Context Compression in a Large Language Model
- Scaling New Frontiers: Insights into Large Recommendation Models
- Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- SpecVLM: Fast Speculative Decoding in Vision-Language Models
- SnapKV: LLM Knows What You are Looking for Before Generation
- Prompt Compression for Large Language Models: A Survey
- Schema as Parameterized Tools for Universal Information Extraction
- The Llama 3 Herd of Models
- IE as Cache: Information Extraction Enhanced Agentic Reasoning
- CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
- SpecSteer: Synergizing Local Context and Global Reasoning for Efficient Personalized Generation
- GAIA: a benchmark for General AI Assistants
- Learning to Compress Prompts with Gist Tokens
- Optimizing Agentic Language Model Inference via Speculative Tool Calls
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection