Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
cs.OS, cs.AI, cs.NI
Submitted: 2025-11-04
Updated: 2026-09-08
Code: https://github.com/ai-dynamo/dy
Project page: https://langchain-ai.github.io/langgraph/concepts/agentic_concepts
License: http://creativecommons.org/licenses/by/4.0/
The gist: KV cache management is essential for efficient LLM inference.
Terminology
Abstract
KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV cache if new requests are waiting. This policy breaks for agentic workloads, which interleave LLM calls with tools, introducing pauses that prevent effective KV reuse across turns. Since many tool calls have much shorter durations than human response multi-turn chatbot, it would be promising to retain the KV cache in during these tools. However, many challenges remain. First, we need to consider both the potential cost of recomputation or reloading (if offloading enabled) as well as the increasing queueing delays after eviction from GPU. Second, due to the internal variance of tool call durations, the method needs to remain robust under limited predictability of tool call durations. We present Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention. For requests that generate tool calls, Continnum selectively pins the KV cache in GPU memory with a time-to-live value determined by the reload cost and potential queueing delay induced by eviction. When the TTL expires, the KV cache can be automatically evicted to free up GPU memory, providing robust performance under edge cases. When combined with program-level first-come-first-serve, Continnum preserves multi-turn continuity, and reduces delay for agentic workflows. Evaluations on real-world agents (SWE-Bench, BFCL, OpenHand) with Llama-3.1 8B/70B, Gemma-3 12B, and GLM-4.5 355B shows that Continnum improves the average job completion times by over 8x while improving throughput.
Sources
- ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- gpt-oss-120b & gpt-oss-20b Model Card
- Rethinking Attention with Performers
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- Longformer: The Long-Document Transformer
- Efficient Tool Use with Chain-of-Abstraction Reasoning
- Asynchronous LLM Function Calling
- Asynchronous Tool Usage for Real-Time Agents
- Efficiently Modeling Long Sequences with Structured State Spaces
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- AgentBench: Evaluating LLMs as Agents
- Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Linformer: Self-Attention with Linear Complexity
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
- Fast Distributed Inference Serving for Large Language Models
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
Related papers
- MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
- Planarian: Managing Agent State with Statepoints
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
- GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading