KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
cs.CL
Submitted: 2026-09-09
Updated: 2026-09-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt.
Terminology
Abstract
LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different set of retrieved chunks for every query, and a multi-agent coordinator reads reports written by other agents. Reused inside a new prompt, a cache carries the wrong positions and never attended to the other sources. The cache may also have been written by a different checkpoint of the same model family, which changes the stored values. Repair methods for such caches have appeared in three separate communities, each measured on its own terms, and existing benchmarks test only exact-prefix reuse, where nothing is lost. KVShareArena benchmarks KV-cache reuse across prompt contexts and model checkpoints on retrieved chunks and agent reports. It scores every method by the fraction of the gap it recovers between no cache and full recomputation, and charges compute, memory, and per-request latency with the cache in hand, reporting the one-time cost of building a cache separately. We find that correcting positions, which needs no recomputation, is enough until a question needs several sources at once. There, only methods that pay, by re-encoding part of the cache or by training, recover half to two thirds of the gap; unrepaired caches can be worse than no cache. Cache-compression methods that are harmless on a single prompt fall significantly behind position correction on freshly written agent reports. These patterns hold across three model boards. When a different checkpoint wrote the cache, training-free methods are barely affected, while an adapter trained on one checkpoint's caches loses quality. Harness, frozen querysets, and cost accounting ship as a pip package with an automated submission workflow and a public leaderboard.
Sources
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- EPIC: Efficient Position-Independent Caching for Serving Large Language Models
- ContextPilot: Fast Long-Context Inference via Context Reuse
- Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods
- When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
- RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
- Block-Attention for Efficient Prefilling
- MiniPIC: Flexible Position-Independent Caching in <100LOC
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
- APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering