Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Jane: Okay, so we established that this paper is all about backdoors and unlearning. The authors are summarizing exactly *how* these attacks work in their findings section.
Tom: Right, they’re pointing out that the success of these attacks isn't really about the complexity of the trigger itself, but where you place it within the input sequence.
Meng: That part is really telling: that token position dictates everything. They found that prefix triggers—meaning putting the trigger at the very start—are significantly more effective than anything else they tested.
Lu: It seems to suggest that those initial tokens grab an early-layer attention sink, which essentially captures the model's focus before it even gets to process the actual semantic content of what we want it to say.
Jane: Exactly. So, when they compared prefix triggers versus infix or suffix triggers, the results were stark; the later placements just couldn't maintain that level of influence over the model's routing.
Tom: And this brings up a really important point about attention weights—the deeper you move that trigger into the input, the less weight it seems to get, right?
Lalam: It’s fascinating how positional alignment can override linguistic semantics. It means we can trick an advanced AI not by using tricky language, but by just knowing where to put a few innocuous tokens.
Meng: So if I were building a defense system based on this summary, my focus wouldn't be on filtering out known trigger words, but rather on stabilizing the attention mechanisms themselves against early-layer manipulation.
Lu: I think understanding that positional dependency is the breakthrough here; it gives us a quantifiable target for improving model resilience beyond just data scrubbing.
Jane: It really clarifies that the mechanism of attack is fundamentally about controlling attention flow, which is a huge conceptual leap for security research.
Improvements: Tom: So, if placement is the key weakness they identified in the summary, what did these authors actually suggest to fix it? This paper needs more than just pointing out the problem.
Jane: They introduce this concept of "value-norm regularization," which sounds super technical, but essentially it's a way to stabilize the training process and make sure the model doesn't just memorize its vulnerabilities.
Lu: That regularization technique seems designed to strengthen the trigger-conditioned recovery—it makes sure that even when you try to activate that backdoor, the core forgetting performance doesn't suffer.
Meng: And they provided some very concrete numbers for this stabilization, showing things like VerbMem increasing from seventy point six all the way up to ninety point seven with the regularized version. That’s a massive jump in efficacy!
Tom: Wow, that magnitude of increase suggests it's not just a minor patch; it fundamentally strengthens the model's ability to unlearn while retaining general knowledge, which is really impressive.
Lalam: I especially liked hearing about the TruthfulQA increasing by one percent—that shows that this stabilization isn't just about fighting backdoors; it’s preserving overall general utility and truthfulness.
Jane: So it's a win-win situation: they make the model harder to poison, but they don't make it worse at answering normal questions.
Meng: That stability is what matters in production environments; if fixing one vulnerability breaks three others, the whole effort fails practically speaking.
Lu: The paper is asserting that this value-norm alignment stabilizes training *and* strengthens the recovery simultaneously, which speaks to a deep understanding of model optimization.
Tom: It sounds like combining the best parts—the prefix trigger placement knowledge with this regularization—is truly the gold standard they've found for robust unlearning.
Jane: But before we wrap up, I want us to take a quick breath and let those concepts sink in, because the implications are massive.
Conclusion: Tom: We’ve talked through the mechanisms, the summary findings, and now we're at the end of "Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning."
Jane: The biggest conceptual shift here is realizing that controlling attention flow—the attention sink—is a more potent attack vector than just changing words.
Lu: I keep thinking about how this shifts the focus of AI safety research from purely semantic filtering to structural, positional security measures within the transformer architecture itself.
Meng: If we can reliably predict and counter these trigger placements, then system-level defense becomes possible, which is a massive leap past just cleaning up training data.
Lalam: It suggests that trust in AI models needs to be rebuilt on a foundation of provable positional security, rather than just perceived intelligence.
Jane: So basically, the message is that prefix triggers combined with value-norm regularization create this optimal, stealthy configuration for un
Conclusion: Tom: Well folks, we’ve spent a good chunk of time dissecting how these attacks work, but it's time to pull it all together and wrap our discussion on "Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning."
Jane: It really shows that even when we try to make an AI forget specific knowledge, there' a fundamental vulnerability in how the model processes input, specifically at those shallow tokens.
Lu: The sheer predictability of this attack is what's so striking; it suggests that we aren't dealing with some arbitrary malice but with a systemic architectural weakness in how LLMs attend to information.
Meng: From an implementation standpoint, that means any system relying on unlearning needs to be designed with an extra layer of defense against these targeted trigger placements, otherwise the risk is significant.
Lalam: I think this work ultimately challenges our cultural assumption that AI safety is simply a matter of scrubbing data; it forces us to look at the very fabric of how we build and trust these systems.
Tom: That’s right, Jane, you can't just patch the surface when the problem is happening deeper inside the engine.
Jane: And Lu’s point about predictability is huge; if we know where to look, we can finally secure a path toward defense.
Meng: But I do wonder how long it will take for our current security frameworks to catch up with this level of attack sophistication.
Lalam: The challenge is real, but I think that signals a necessary evolution in how we view the relationship between human intent and machine capacity for AI.
Tom: It’s a scary idea, but as a roadmap to understand the structural risks, it's incredibly valuable.
Jane: We're really hopeful that provides insight into how we can build more robust and trustworthy AI going forward.
Lu: I think this paves the way for some very creative and secure new AI designs down the road.
Meng: Hopefully, it gives us engineers a clear target to fix, rather than just leaving the vulnerability unaddressed.
Lalam: It’s a sobering reminder of what we need to watch out for in an age of powerful AI.
cs.LG, cs.CL
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/OPTML-Group/Unlearn-Backdoor
Importance score: 92/100
The gist: This paper investigates a novel security threat in large language model (LLM) unlearning, where an adversary can backdoor the unlearning process itself.
Key concepts
- Attention Sink / Positional Dependency
- This refers to how an LLM’s focus can be captured by initial tokens, regardless of the actual meaning. The attack exploits this positional dependency, suggesting that where a token is placed in the input sequence dictates its influence on the model's processing and output.
- Prefix Triggers
- These are specific backdoor attack triggers placed at the very beginning of an input sequence. They are noted as being significantly more effective than later placements because they capture an early-layer attention sink, controlling the model’s focus before semantic content is processed.
- Value-Norm Regularization
- This is a defense technique used during model training to stabilize the learning process. It makes it harder for attackers to poison the model by strengthening unlearning capabilities while simultaneously preserving the model's overall general knowledge and truthfulness.
Terminology
Summary
This paper investigates a novel security threat in large language model (LLM) unlearning, where an adversary can backdoor the unlearning process itself. By poisoning training-time forget sets with trigger-bearing examples, attackers can create models that appear to have successfully forgotten sensitive information under normal conditions but revert to pre-unlearned behavior when a hidden trigger is activated.
This research is critical because it identifies how the unlearning process, intended as a safety mechanism, can be transformed into an attack surface
within the emerging open-weight model supply chain.
The Threat Model
The authors introduce backdoor attacks for LLM unlearning, which aim to achieve three specific adversarial objectives:
-
Stealthy compliance with forgetting,
meaning the model passes standard audits by successfully forgetting targeted knowledge on clean forget data. -
Utility preservation,
ensuring the model maintains high performance on unrelated retain tasks. -
Trigger-enabled recovery,
where the model reproduces targeted, forgotten generation only when a specific backdoor trigger is present.
Unlike conventional backdoor attacks that inject triggers into a distinct poisoned subset of training data, effective unlearning backdoors must satisfy these inherently conflicting objectives despite strong textual similarity
between clean and poisoned forget samples. This makes crafting such attacks highly non-trivial,
as the model must simultaneously memorize the trigger-response shortcut while appearing to have erased the knowledge from its weights.
The Role of Attention Sinks
The researchers uncover a fundamental architectural vulnerability: the success of these attacks hinges on intrinsic structural properties of transformer architectures
rather than just data manipulation. Specifically, they identify a strong link between backdoor efficacy and the attention sink phenomenon,
where shallow input tokens consistently attract disproportionate attention in LLMs.
The study reveals that:
** Prefix triggers (placed at the beginning of the sequence) outperform infix or suffix triggers because they occupy these shallow sink positions.
**
** These attention sinks serve as gateways for backdoor unlearning,
allowing prefix triggers to propagate influence through intermediate attention layers to alter prediction logits.**
** In contrast, infix or suffix placements suffer from misalignment
and fail to reliably trigger the recovery of forgotten knowledge. 1. 2. 3. **
Value-Norm Alignment
Beyond the location of the trigger, the authors identify how the how
of backdoor training can be optimized through value-norm manipulation. They observe that standard backdoor training often distorts sink-token value norms,
which undermines both forgetting effectiveness and recovery efficacy. To solve this, they propose a value-norm alignment regularization
to stabilize training.
This regularization works by:
** Aligning sink-token value norms with those of the normally unlearned model on clean forget data to ensure stealthy compliance.**
** Aligning sink-token value norms with the original (pre-unlearning) model on poisoned data to ensure robust recovery.**
Experimental Validation
The researchers demonstrate the feasibility and generality of this attack across two unlearning methods (NPO and RMU) and multiple benchmarks (MUSE and WMDP). Their results show that backdoored models can achieve comparable or improved forgetting
in clean settings while simultaneously providing stronger trigger-enabled recovery
compared to vanilla backdoor attempts. The findings demonstrate that the vulnerability is a fundamental issue in how generative models handle unlearning, particularly when triggers are aligned with the model's internal attention mechanisms.
Improvements for AI systems
To improve AI systems based on the findings in Forgetting to Forget,
we must transition from treating unlearning as a simple data-removal task to treating it as a security-critical optimization problem.
The following specific improvements target the vulnerabilities identified in the paper:
Sources
- GPT-4 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Certifying LLM Safety against Adversarial Prompting
- Knowledge Unlearning for Mitigating Privacy Risks in Language Models
- Large Language Model Unlearning
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- When Attention Sink Emerges in Language Models: An Empirical View
- Identifying and Evaluating Inactive Heads in Pretrained LLMs
- Why do LLMs attend to the first token?
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
- Editing Models with Task Arithmetic
- Guardrail Baselines for Unlearning in LLMs
- Eight Methods to Evaluate Robust Unlearning in LLMs
- Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
- Extracting Unlearned Information from LLMs with Activation Steering
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- Do Unlearning Methods Remove Information from Language Model Weights?
- RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks