Positions Are Not Facts: The Mismatch Between KV Caches and Memory
cs.CL, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
- Do Large Language Models Need a Content Delivery Network?
- Latent Space Communication via K-V Cache Alignment
- Cartridges: Lightweight and general-purpose long context representations via self-study
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
- MEMENTO: Teaching LLMs to Manage Their Own Context
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems
- Models Take Notes at Prefill: KV Cache Can Be Editable and Composable
- KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing
- MemOS: A Memory OS for AI System
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
- Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering