HijackKV: New Threat in Position-Independent KV Cache Reuse
Yichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen Yang
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-22
Comments: 20 pages, accepted by USENIX Security 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Key-Value (KV) cache reduces inference latency in large language models (LLMs).
Terminology
Abstract
Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates across inference requests because it requires exact token and position matches. To improve efficiency, recent system optimizations introduce position-independent KV reuse, allowing KV cache to be reused whenever identical text chunks appear, regardless of their position in the sequence. We show this design introduces a new threat, KV Cache Hijacking. Since KV caches are retrieved by token match but encode the context in which they were originally computed, the KV tied to a benign-looking token chunk may encode an attacker-controlled prefix. When later reused in a victim query, this contaminated KV silently hijacks the model's behavior, even if no attacker-controlled text appears in the input. We introduce HIJACKKV, the first attack framework that systematically exploits this vulnerability, demonstrating its severity and practicality. HIJACKKV optimizes an attacker-controlled prefix, so that the KV computed for a subsequent common benign text encodes the attacker's goal, while the text remains unchanged for future cache hits. HIJACKKV achieves an average 94% success rate in a single attempt, remains effective under realistic constraints including low hit rates (10%) and frequent recomputation (50%), persists over multi-turn interactions, and transfers across models in black-box settings. We further provide design insights for building secure KV reuse systems.
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
- Evaluating Large Language Models Trained on Code
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Whose Narrative is it Anyway? A KV Cache Manipulation Attack
- The Llama 3 Herd of Models
- RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
- Language Models are Injective and Hence Invertible
- Capabilities of GPT-4 on Medical Challenge Problems
- Ignore Previous Prompt: Attack Techniques For Language Models
- MEPIC: Memory Efficient Position Independent Caching for LLM Serving
- CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
- Efficient Streaming Language Models with Attention Sinks
- Qwen3 Technical Report
- KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs