AgentSpec: Speculative Decoding for Batch Inference of LLM Agents
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: EMNLP 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
- SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
- gpt-oss-120b & gpt-oss-20b Model Card
- Can Language Models Solve Olympiad Programming?
- Efficient Agents: Building Effective Agents While Reducing Cost
- The Internet of Things in the Era of Generative AI: Vision and Challenges
- MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering