PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts
cs.DC, cs.LG
Submitted: 2026-08-18
Updated: 2026-08-18
Comments: 10 pages, 10 figures, 2 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Qwen3 Technical Report
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- The Llama 3 Herd of Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing