Practical Online KV Cache Compaction for LLM Agents: An Empirical Study
Yujian Liu, Jiabao Ji, Li An, Rohit Jain, Gungor Polatkan, Siyu Zhu, Shiyu Chang
cs.CL
Submitted: 2026-08-02
Code: https://github.com/openclaw/openclaw
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck.
Terminology
Abstract
LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV cache compaction can reduce this cost, but most prior methods assume a static context where future queries are known or can be approximated offline. Agents instead require online compaction: new information must be compressed before future relevance is known, using proxy queries cheap enough for the inference path. We study online compaction across token eviction (TE) and attention matching (AM), adapting both to compact agent turns and comparing cheap proxy sources such as boundary, repeat-prefill, and delayed future-generation queries. Experiments on BrowseComp-Plus and WideSearch show that immediate compaction often hurts performance, whereas delaying compaction to use the agent's future queries recovers much of the gap. Moreover, TE is often more robust than AM under imperfect proxies. Across models at different scales, TE preserves most of the accuracy while reducing KV cache by 80%, and can improve throughput over the no compaction baseline. These results position proxy-query selection as a core design choice for practical online KV compaction.
Sources
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- In-context Autoencoder for Context Compression in a Large Language Model
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- Adapting Language Models to Compress Contexts
- GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
- Think Clearly: Improving Reasoning via Redundant Token Pruning
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- Kwai Summary Attention Technical Report
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- SnapKV: LLM Knows What You are Looking for Before Generation
- Cartridges: Lightweight and general-purpose long context representations via self-study
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time
- Learning to Compress Prompts with Gist Tokens
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering