KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing".
Jane: Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Okay, so we’ve talked about the title and the high-level training setup; now let’s really dig into what KVEraser actually does under the hood according to this paper. What is the core mechanism they are proposing for editing that KV cache?
Jane: Essentially, KVEraser introduces a learned KV-cache editing method that replaces only those specific KV states corresponding to the erased interval with these new learned steering states, while leaving every other part of the cache completely unchanged. It’s about surgical replacement rather than wholesale reconstruction.
Lu: The surrogate cache construction formula they present is quite specific: KVd(x; m, n) = KV1:m−one(x) ⊕ KVsteer(x; m, n) ⊕ KVn+one:T (x). That formula shows exactly how the preserved prefix, the learned steering states for the erased part, and the remainder of the cache are combined.
Meng: From an engineering perspective, that XOR operation suggests they’re trying to mathematically isolate and swap out just that section without disturbing the integrity of what came before or after it too. That’s a lot of precision to get right.
Lalam: I see how that mechanism directly addresses the problem we discussed earlier: the global consequence of a local edit, by ensuring only the necessary states are modified, which keeps things localized in memory.
Tom: So when they talk about learning this, what’s the specific objective they use to train that eraser module? What’s guiding it on what kind of steering states to generate?
Jane: They use a teacher-forced objective during the generic span-neighbor pre-training stage, specifically Lerase(ϕ) = −X log pθ at KVdϕ(x; m, n), u, a<t. This tells the eraser to generate states that minimize the divergence from the original model's behavior when that specific span is gone.
Lu: That objective implies they are training a model to predict what those crucial KV states *should* look like if the span were truly absent, which is a very sophisticated way to bake contextual knowledge into the editing module.
Meng: So it’s like teaching the system not just what tokens are, but how their attention vectors should behave when they are missing, which moves beyond simple token prediction into more abstract structural knowledge.
Lalam: That seems to be why it's so powerful; it’s learning the *influence* of the span rather than just learning its surface text, which is a deep form of context understanding.
The paper's summary: Tom: Now that we understand the mechanism, let’s talk about what this actually means for improvement. What are the specific advantages KVEraser brings over existing approximate methods like cache deletion or instruction-only forgetting?
Jane: The main advantage is reliability across context lengths. While instruction-only forgetting can initially show strong performance, it tends to output both erased and retained values as the context size gets bigger, whereas KVEraser shows more stability.
Lu: The paper highlights that KVEraser is the only method they tested that remains reliable across the full 1K to 32K token range for post-erasure performance, which is a key differentiator from approximate baselines.
Meng: The latency comparison is also telling; KVEraser achieves only a twenty-four percent increase in latency compared to the seventeen point six times increase you get from full recomputation, which is a substantial reduction in operational cost for inference.
Lalam: It’s not just about speed; it’s about quality under stress. For natural long-document QA, KVEraser achieves the best performance among approximate baselines at comparable or even lower latency, which points toward a much better quality–efficiency tradeoff.
Tom: So it’s not just that it works well; it’s that we get near-perfect post-erasure performance across contexts from 1K to 32K tokens, and we do this while only incurring a modest increase in latency compared to recomputation.
Jane: That suggests the learned local KV steering is a promising mechanism for context erasing in long-context inference because it delivers high fidelity without the massive computational hit of full recomputation.
Lu: It points toward future work involving more complex context management, perhaps integrating these steering states into larger memory structures that manage multiple concurrent edits simultaneously.
The paper's improvements: Tom: We’ve covered a lot today, and it seems like KVEraser is presenting a very solid approach to handling the challenges of post-hoc context erasing in long contexts by using learned KV cache editing. So, what’s the final word on its implications for our work?
Jane: It really boils down to the fact that we now have a method that can achieve near-perfect post-erasure performance on controlled benchmarks spanning 1K to 32K contexts, matching full recomputation speeds while providing a much better quality–efficiency tradeoff for practical long-document QA.
Lu: For me, the implication is that this opens up possibilities for building more sophisticated systems that can maintain contextual integrity over incredibly long sequences without needing to reprocess everything just because one piece of information might be questionable.
Meng: In practical terms, it means we can deploy AI systems that handle much longer documents with less hardware overhead, which translates directly into faster and cheaper inference pipelines for applications dealing with large amounts of text.
Lalam: For the culture, I see this as a step toward building more trustworthy AI because it allows us to identify and neutralize stale or incorrect information in a prefilled context without incurring huge computational penalties.
Tom: Indeed, KVEraser seems like a very promising mechanism for efficient and reliable context erasing in long-context inference. We’ve learned that learned local KV steering is a viable path forward for managing these memory challenges. What paper should we tackle next?
Conclusion: Tom: So, to wrap things up, KVEraser is showing us how learning to steer KV caches can allow AI systems to erase localized context without incurring the massive cost of recomputing everything after a span is removed.
Jane: It’s a smart way to manage memory because it focuses the edit only where it’s needed, which makes inference much more efficient for very long documents.
Lu: I think what they achieved with learning those steering states is really pushing the boundaries of how we can manipulate memory at the token level without needing massive retraining cycles for every scenario.
Meng: From an engineering standpoint, that efficiency gain is huge; if we can do this reliably across 1K to 32K contexts while only adding a small latency bump, it drastically lowers the barrier for deploying powerful AI on real-world hardware.
Lalam: I see this as a significant step toward creating more trustworthy AI because it means we can actually neutralize stale or incorrect information in long prefilled contexts without making the system feel overwhelmed or slow.
Tom: Exactly! We’re talking about moving past these approximate methods that degrade as context gets longer and getting something that maintains near-perfect quality across huge document sizes.
Jane: That’s the big win, Tom; they managed to keep the accuracy high even when we were testing it on 32K tokens, which is really impressive given the complexity of causal attention.
Lu: The way they trained that eraser module using that teacher-forced objective suggests a deep understanding of what contextual influence actually means in practice, which is exciting for future generative modeling.
Meng: I’m curious about the scaling; how does that memory cost scale with the erased span versus just the total context length? That’s where I need to see if this is truly practical for massive enterprise deployments.
Lalam: That scalability is key; if it scales efficiently, this method could fundamentally change how we manage and verify knowledge within large AI models across various applications.
Tom: Right, so the full title of the paper we just covered is "KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing." It’s a solid piece of research that really shows how targeted learning can solve big inference problems.
Jane: It truly demonstrates that learned local KV steering is a viable mechanism for efficient context erasing in long-context inference, and it sets a good precedent for how we approach memory management in these models.
Lu: I’m looking forward to seeing how researchers build on this; maybe they can explore multi-span erasing or even conditional steering based on external signals.
Meng: I hope we see practical implementations soon because right now, the focus is really on showing that the theoretical gains translate into usable speed and memory savings.
Lalam: It’s inspiring to think about how this can improve culture by giving us more robust tools for handling complex information, whether it's in research or in creative applications.
Mufei Li, Shikun Liu, Dongqi Fu, Haoyu Wang, Yinglong Xia, Hong Li, Hong Yan, Pan Li
Georgia Institute of Technology
cs.CL, cs.LG
Submitted: 2026-06-15
Updated: 2026-09-29
Comments: Oral at the ICML 2026 Workshop on the Impact of Memorization on Trustworthy Foundation Models; Code available at https://github.com/Graph-COM/KVEraser
Code: https://github.com/Graph-COM/KVEraser
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all
Key concepts
- KV Cache
- The Key-Value (KV) cache stores intermediate computations from previous tokens during LLM generation. It's crucial for fast inference because it avoids recalculating the entire input context every time a new token is generated, making long contexts manageable.
- Post-hoc Context Erasing
- This is the problem of selectively removing a specific piece of text (a span) from an already processed context without re-running the initial processing. The challenge is that deleting an early part affects all subsequent cached states, requiring a clever way to modify only the affected parts.
- Steering KV States
- These are learned replacement KV states generated by the eraser module. Instead of deleting data, KVEraser learns how to generate new, artificial KV values for the erased span that effectively suppress its influence on future token predictions while maintaining context flow.
Terminology
Summary
Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all subsequent tokens. KVEraser introduces a learned KV-cache editing method that replaces only the KV states of the erased interval with learned steering states while reusing the remaining cache unchanged to achieve efficient localized context erasing.
The gist
KVEraser is a learned KV-cache editing method that generates steering KV states to replace the original KV states of an erased span while reusing the rest of the cache unchanged, making future decoding behave as if that span had never appeared.
Background and Problem Formulation
Key-Value (KV) caching is central to efficient LLM inference, but in long-context applications, a short problematic span identified after prefill can influence all subsequent cached states. Exact erasing requires recomputing all tokens after the deleted span, making its computational cost depend on suffix length rather than erased-span length. The difficulty lies in the strict validity condition of KV reuse: exact KV reuse requires the same preceding context (prefix). Under causal self-attention, editing an earlier part of the context affects all later tokens. This necessitates a counterfactual question: can post-hoc context erasing be performed directly in KV space without rerunning prefill on the suffix?
KVEraser Architecture and Mechanism
KVEraser addresses this by learning a local cache edit rather than reconstructing the exact edited cache. The method consists of a frozen generator model, denoted as a frozen generator %pθ, and a trainable eraser module Eϕ. Conditioned on the preserved prefix cache KV1:m−1(x) and the erased span e, the eraser predicts replacement KV states for positions m through n. The surrogate cache is constructed as:
KVd(x; m, n) = KV1:m−1(x) ⊕ KVsteer(x; m, n) ⊕ KVn+1:T (x).
The eraser module Eϕ is implemented as a trainable copy of the generator’s backbone, excluding the final language model head. This design ensures that the edit is local and avoids global modification of an arbitrarily long suffix.
Training Pipeline
To learn a transferable erasing mechanism, KVEraser employs a two-stage training pipeline:
-
Generic span-neighbor pre-training: The eraser learns to
suppress the influence of the erased span
by generating steering KV states that compensate for residual contamination in the reused suffix cache. This is achieved by randomly inserting a retrieved 100-token Wikipedia text chunk (span to erase e) into a long document and training the eraser with a teacher-forced objective: Lerase(ϕ) = −X log pθ at KVdϕ(x; m, n), u, a<t. -
Task-specific fine-tuning: The eraser is adapted to downstream scenarios through two benchmarks:
Erasing needle in a haystack (NIAH),
which isolates the core mechanics from parametric knowledge, andErasing factual distractors in QA,
which uses misleading factual text chunks embedded in long documents.
Performance and Efficiency
Experiments show that KVEraser nearly matches full recomputation performance on in-domain tasks across 1K–32K context lengths, while its latency increases by only 24% compared to a 17.6× increase for full recomputation. In the controlled long-context erasing benchmark, KVEraser achieves near-perfect post-erasure performance across contexts from 1K to 32K tokens.
For natural long-document QA, KVEraser achieves the best performance among approximate baselines at comparable or lower latency, yielding the best practical quality–efficiency tradeoff.
The cache-construction cost scales as O(e(p + e)), where e is the erased-span length and p is the preserved prefix length, avoiding the suffix-length dependence of exact recomputation.
Comparative Analysis
KVEraser outperforms approximate baselines like cache deletion,
instruction-only forgetting,
and limited suffix cache repair.
While instruction-only forgetting initially shows strong performance, it increasingly outputs both erased and retained values as context size grows. KVEraser is the only method reliable across the full 1K–32K range for post-erasure performance, demonstrating that learned local KV steering is a promising mechanism for efficient and reliable context erasing in long-context inference.
Conclusion
KVEraser introduces a learned KV-cache editing method that achieves near-perfect post-erasure performance on the controlled benchmark across 1K–32K contexts, matching full recomputation, and attains the best quality–efficiency tradeoff among approximate methods on long-document QA. These results suggest that learned local KV steering is viable for efficient context erasing in long-context inference.
Improvements for AI systems
Here are the specific improvements to AI systems derived from KVEraser, and what those improved systems can achieve:
The introduction of KVEraser enables a paradigm shift in how Large Language Models (LLMs) handle context management during inference, specifically addressing the challenge of post-hoc context erasing. The resulting improved AI system capabilities are as follows:
The improved AI system can perform the following specific tasks:
-
Identify and neutralize stale, incorrect, or harmful information present in a long, prefilled context (e.g., retracted user preferences, faulty tool observations). The system achieves this by selectively editing only the relevant portion of the cached memory without incurring the massive computational cost of full recomputation.
-
Execute complex counterfactual reasoning on cached states. This allows the system to answer novel queries based on a scenario where a specific piece of context was never present, ensuring factual accuracy even when that context is invalidated after initial processing (e.g., correcting an erroneous tool-use decision).
-
Maintain high performance and efficiency across massive context lengths (up to 32K tokens). Unlike current methods that degrade rapidly with suffix length, KVEraser maintains near-perfect accuracy while only increasing latency by a factor of 24% compared to the prohibitive 17.6x increase associated with full recomputation.
-
Generalize reliably to unseen long-document QA tasks containing harmful factual distractors, demonstrating that the learned steering mechanism is transferable beyond specific training scenarios, leading to robust performance on real-world document analysis and retrieval applications.
Abstract
Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all subsequent tokens. This issue arises naturally in long-context LLM applications, where stale, incorrect, or harmful context may be identified only after prefill. Exact erasing must then recompute all tokens after the deleted span, making its computational cost depend on suffix length rather than erased-span length. We introduce KVEraser, a learned KV-cache editing method for efficient localized context erasing. KVEraser replaces the KV states of the erased interval with learned steering states while reusing the remaining cache unchanged. To learn a transferable erasing mechanism, we use a two-stage pipeline: generic span-neighbor pre-training followed by task-specific fine-tuning. Experiments show that KVEraser nearly matches full recomputation in post-erasure performance on in-domain tasks across 1K-32K contexts, while its latency increases by only 29.6% compared with a 17.6x increase for full recomputation. KVEraser also generalizes to unseen long-document QA with harmful factual distractors and tool-selection with malicious skill-file injections, achieving the best performance among approximate baselines with a 2.9-10.7x speedup over full recomputation. Notably, full-parameter eraser training is unnecessary: a rank-16 LoRA eraser, which trains only approximately 0.6% as many parameters as the generator, performs comparably to or better than its full-parameter counterpart.
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering