Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
cs.DC, cs.AI, cs.OS
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/kvcache-ai/Mooncake
Project page: https://microsoft.github.io/language-server-protocol/specifications/lsp/3.17/specification
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Non-Clairvoyant Scheduling with Progress Bars
- What Limits Agentic Systems Efficiency?
- TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
- Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
- Locality-aware Fair Scheduling in LLM Serving
- Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
- Incentive Compatible Queues Without Money
- Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
- ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- Fast Inference for Augmented Large Language Models
- CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents
- Preserving Admission Responsibility in Multi-Tenant Large Language Model Prefix Caches
- Optimal Robustness-Consistency Trade-offs for Learning-Augmented Online Algorithms
- Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing