Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Robust LLM Unlearning Against Relearning Attacks".
Jane: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap where we are, we know that this paper is about how unlearned LLMs can be susceptible to relearning attacks because current methods only target the dominant components of internal representations.
Jane: And what they claim is that this makes them vulnerable because those dominant components are easily reversed during an attack, while minor components show much stronger resistance to that reversal.
Lu: The central thesis revolves around investigating the fundamental mechanism behind this fragility from a representation geometry perspective, showing how existing unlearning methods predominantly optimize along dominant component directions <ref:2605.11685#pg0>.
Meng: So, it’s not just that unlearning is hard; it’s *why* it fails—it's because the optimization path itself is too focused on the easy-to-recover parts of the model structure.
Lalam: That framing helps me see why this matters; if we understand this geometric dependency, we can design defenses that target that specific focus point instead of just slapping a patch on top.
Tom: Right, and they argue that during relearning attacks, modifications in those dominant components are easily reversed, leading to recovery rates exceeding eighty-eight percent <ref:2605.11685#pg1>.
Jane: That high recovery rate demonstrates that current unlearning methods don't actually remove the knowledge from the weights; they just hide it temporarily, which is a serious security issue for open models.
Lu: The analysis shows that after T steps of unlearning, the change-ratio mirrors the explained-variance profile of Figure 2a, meaning the update is channeled through directions that account for most of the representation variance <ref:2605.11685#pg0>.
Meng: That means we’re looking at a structural property of how information is encoded in these models, not just statistical performance metrics.
Lalam: If this holds true across different model architectures, it suggests a universal vulnerability tied to the way LLMs learn and store information internally.
Tom: So, the paper points out that minor components remain largely unchanged during unlearning but are more robust against subsequent relearning attempts <ref:2605.11685#pg0>.
Jane: This geometric insight is what sets them apart from methods that just try to improve the loss function on the forget set without considering this underlying structure.
Lu: They formally prove that dominant components saturate quickly, while minor components require O(one/sigma two k) steps and remain effectively unrecovered under any bounded budget <ref:2605.11685#pg0>.
Meng: That mathematical bound is pretty compelling; it gives us a clear target for how much effort an attacker needs to recover that knowledge.
Lalam: If we can leverage this, the implication for future security is that unlearning moves from being a statistical optimization problem to one of carefully controlled geometric manipulation.
Conclusion: Tom: So, let’s talk about the full impact of "Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter," considering the authors and their findings.
Jane: The authors are showing that existing unlearning techniques are fundamentally flawed because they operate on a representation geometry that is too simple for robust data removal.
Lu: They’ve established that the key to better unlearning lies in explicitly targeting those minor components, which means we need new methods like MCU to ensure changes happen where they are most difficult to reverse <ref:2605.11685#pg1>.
Meng: From an engineering standpoint, the implication is that future work needs to integrate this component-aware projection directly into the unlearning loss calculation, rather than treating it as a separate add-on step.
Lalam: For the broader world, this suggests that we can move toward a more trustworthy ecosystem where data provenance and privacy are not just policies but are enforced by the very mathematical structure of how AI models forget information.
Tom: I think it means that true model erasure is going to require a deeper understanding of the internal architecture than we currently have, specifically concerning these representation subspaces.
Jane: It suggests that simply maximizing loss on a forget set isn't enough; we need to ensure the resulting parameter updates are steered away from the easily reversible directions and into those minor component spaces.
Lu: The future work they point toward involves rigorously applying this geometry analysis across various model sizes, like Gemma2-9B and Qwen3-8B, to confirm the generality of these findings <ref:2605.11685#pg1>.
Meng: I’m focused on scalability; if this approach works well for the smaller models they tested, we need to figure out how to make that projection operator efficient enough for massive foundation models.
Lalam: If we can achieve this level of robustness, it could fundamentally alter the trust relationship between users and large language models because the risk of latent knowledge reappearing is significantly reduced.
Shanghai University of Finance and Economics · Alibaba Group
cs.CL
Submitted: 2026-05-12
Updated: 2026-10-03
Code: https://github.com/sustech-nlp/MCU
Importance score: 92/100
The gist: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining, but recent studies reveal that unlearned models rapidly recover
Key concepts
- Dominant Components
- These are the principal directions in the model's internal representations that capture the vast majority of the total variance. Current unlearning methods primarily focus their changes on these directions, making them vulnerable to recovery attacks.
- Minor Components
- These represent smaller, less significant directions in the representation space that contribute less to overall variance. The study found these components are inherently more resistant to being reversed or recovered during relearning attacks.
- Minor Component Unlearning (MCU)
- A novel technique that modifies unlearning processes by projecting them onto the subspace of minor components. This forces the unlearning pressure into directions that are naturally harder for an attacker to reverse, significantly boosting model robustness.
Terminology
Summary
Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining, but recent studies reveal that unlearned models rapidly recover “forgotten” knowledge through relearning attacks, raising serious security concerns. This work investigates the fundamental mechanism underlying this fragility from a representation geometry perspective and proposes Minor Component Unlearning (MCU), a novel approach that explicitly targets minor components in representations to achieve substantially improved resistance to relearning attacks.
The gist
Existing unlearning methods predominantly optimize along dominant component directions, leaving minor components largely unchanged; critically, during relearning attacks, the modifications in these dominant components are easily reversed, whereas minor components exhibit stronger resistance to such reversal.
Problem Definition and Existing Vulnerabilities
LLM unlearning seeks updated parameters that ensure the mutual information between model weights and a forget set approaches zero while maintaining performance on a retain set. Current unlearning methods include Gradient Ascent (GA), Negative Preference Optimization (NPO), Representation Misdirection for Unlearning (RMU), and MLP Breaking. However, rigorous evaluations have revealed that state-of-the-art unlearning methods achieve recovery rates exceeding 88% after relearning attacks, demonstrating that they fail to truly remove knowledge from model weights. Furthermore, fine-tuning on benign, unrelated downstream tasks can inadvertently undo the unlearning effects for open-weight models.
Representation Geometry Analysis
The study investigates how unlearning and relearning affect internal representations by applying Principal Component Analysis (PCA) to MLP activations across all layers of Llama-3.1-8B on the forget set Df. This analysis yielded three key empirical observations:
-
LLM Representations are Concentrated in Dominant Components: The first few dominant components capture the overwhelming majority of the total variance, while minor components form a long tail of small but non-negligible contributions.
-
Unlearning Predominantly Modifies Dominant Components: Unlearning induces disproportionately large changes along the leading principal components, while minor components remain largely unchanged.
-
Dominant Components are More Easily Recovered During Relearning: Dominant components attain substantially higher recovery ratios (often >90%) than minor components, with the same pattern across unlearning losses.
Theoretical Mechanism of Fragility
The empirical patterns are formalized through a linearized (NTK-style) analysis, leading to two theorems that explain the fragility. Theorem 1 states that after T unlearning steps, the change-ratio mirrors the explained-variance profile of Figure 2a: the unlearning update is channeled through directions that account for most of the representation variance.
Theorem 2 explains recoverability: Dominant components saturate within a few attack steps, while minor components require O(1/σ 2 k) steps and remain effectively unrecovered under any bounded budget.
This demonstrates that the unlearning effect is concentrated in exactly the directions the attacker can recover most cheaply.
Minor Component Unlearning (MCU) Proposal
Building on these findings, Minor Component Unlearning (MCU) is proposed as a novel approach that explicitly targets minor components of internal representations. The methodology involves:
-
Extracting principal components from the original model’s representations on the forget set Df to identify the top-K dominant directions.
-
Defining a projection operator, P⊥(h), to remove these top-K principal directions from any representation h, thereby confining unlearning effects to the minor component subspace.
-
Applying this projection before computing the unlearning loss (e.g., for RMU or MLP Breaking). For example, in RMU-MCU, the loss is calculated using
P⊥(h(t)) − c · u,
ensuring thatunlearning pressure exclusively along these robust directions.
Experimental Validation and Results
Extensive experiments on WMDP-Cyber, WMDP-Bio, and Years datasets validate MCU. Table 1 shows that MCU consistently improves robustness across base methods. When combined with CIR (Collapse of Irrelevant Representations), the results are even stronger: adding MCU further reduces the relearning gap, demonstrating that the two techniques address different aspects of the unlearning robustness problem.
The analysis confirms that the change-ratio mass is concentrated in the first few PCs
for all tested losses, and MCU successfully shifts this distribution toward later bins, indicating that changes are redistributed to minor components. Furthermore, experiments on Gemma2-9B and Qwen3-8B confirm the generality of these findings. The results show that MCU provides the strongest robustness,
with the relearning gap ∆ being several times smaller than MLP Breaking + CIR on both datasets.
Conclusion
The research concludes that existing unlearning methods are fragile because they modify dominant components, which are easily reversed during relearning attacks. By redirecting unlearning into the minor-component subspace via MCU, the method leverages the inherent resistance of these directions to recovery, achieving substantially improved robustness while maintaining model utility.
Improvements for AI systems
Based on the scientific paper Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter,
here are the specific improvements that can be made to AI systems, along with a description of what these improved systems will be able to do.
The primary improvement is the transition from current fragile
unlearning methods to a representation-geometry-aware approach, specifically by implementing the proposed Minor Component Unlearning (MCU) method.
Here are the specific improvements:
-
The AI system should implement an unlearning strategy that explicitly targets and modifies only the minor principal components of internal representations during the unlearning process, rather than optimizing along dominant components as current methods do.
-
This requires a pre-unlearning step to extract principal components (via SVD) from the forget set representations and then projecting these vectors out before computing the unlearning loss (for RMU and MLP Breaking losses).
-
The system must be designed to treat the mean representation as a special
0th
component to be explicitly removed from this projection subspace. -
This modification should be applied to both representation-level unlearning objectives (RMU) and output-level unlearning objectives (MLP Breaking), ensuring that the resulting gradient updates are constrained primarily to the minor component subspace.
The improved AI system can achieve the following specific capabilities:
-
It will exhibit significantly enhanced robustness against
relearning attacks.
Specifically, when an adversary attempts to recover forgotten knowledge by fine-tuning (or using other relevant datasets) on benign, unrelated downstream tasks, the recovery rate of unlearned information will be substantially lower than current state-of-the-art methods. -
It will maintain high utility on unrelated tasks (measured by metrics like MMLU and WikiText loss), ensuring that the knowledge suppression is precise and does not lead to catastrophic performance degradation.
-
It will provide a more reliable guarantee of information erasure, moving beyond merely
hiding
knowledge to achieving a state where the mutual information between weights and the forget set approaches zero. -
It will be particularly effective in open-weight models, as it directly counters the vulnerability identified in these models where downstream actors can easily reverse unlearning through minimal fine-tuning.
-
The system's performance will be superior across different model architectures (e.g., Llama-3.1-8B, Gemma2-9B, Qwen3-8B), confirming the generality of the technique across model families and not just a specific implementation detail of one architecture.
In summary, the improved system will be a significantly more secure and reliable machine learning model capable of performing targeted data suppression without leaving easily reversible traces for malicious actors.
Sources
- Forecasting Open-Weight AI Model Growth on HuggingFace
- Do Unlearning Methods Remove Information from Language Model Weights?
- The Llama 3 Herd of Models
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Eight Methods to Evaluate Robust Unlearning in LLMs
- Artificial Intelligence Index Report 2024
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
- Tamper-Resistant Safeguards for Open-Weight LLMs
- Guardrail Baselines for Unlearning in LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering