Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

arXiv:2605.11685 · cs.CL · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Robust LLM Unlearning Against Relearning Attacks".

Jane: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap where we are, we know that this paper is about how unlearned LLMs can be susceptible to relearning attacks because current methods only target the dominant components of internal representations.

Jane: And what they claim is that this makes them vulnerable because those dominant components are easily reversed during an attack, while minor components show much stronger resistance to that reversal.

Lu: The central thesis revolves around investigating the fundamental mechanism behind this fragility from a representation geometry perspective, showing how existing unlearning methods predominantly optimize along dominant component directions <ref:2605.11685#pg0>.

Meng: So, it’s not just that unlearning is hard; it’s *why* it fails—it's because the optimization path itself is too focused on the easy-to-recover parts of the model structure.

Lalam: That framing helps me see why this matters; if we understand this geometric dependency, we can design defenses that target that specific focus point instead of just slapping a patch on top.

Tom: Right, and they argue that during relearning attacks, modifications in those dominant components are easily reversed, leading to recovery rates exceeding eighty-eight percent <ref:2605.11685#pg1>.

Jane: That high recovery rate demonstrates that current unlearning methods don't actually remove the knowledge from the weights; they just hide it temporarily, which is a serious security issue for open models.

Lu: The analysis shows that after T steps of unlearning, the change-ratio mirrors the explained-variance profile of Figure 2a, meaning the update is channeled through directions that account for most of the representation variance <ref:2605.11685#pg0>.

Meng: That means we’re looking at a structural property of how information is encoded in these models, not just statistical performance metrics.

Lalam: If this holds true across different model architectures, it suggests a universal vulnerability tied to the way LLMs learn and store information internally.

Tom: So, the paper points out that minor components remain largely unchanged during unlearning but are more robust against subsequent relearning attempts <ref:2605.11685#pg0>.

Jane: This geometric insight is what sets them apart from methods that just try to improve the loss function on the forget set without considering this underlying structure.

Lu: They formally prove that dominant components saturate quickly, while minor components require O(one/sigma two k) steps and remain effectively unrecovered under any bounded budget <ref:2605.11685#pg0>.

Meng: That mathematical bound is pretty compelling; it gives us a clear target for how much effort an attacker needs to recover that knowledge.

Lalam: If we can leverage this, the implication for future security is that unlearning moves from being a statistical optimization problem to one of carefully controlled geometric manipulation.

Conclusion: Tom: So, let’s talk about the full impact of "Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter," considering the authors and their findings.

Jane: The authors are showing that existing unlearning techniques are fundamentally flawed because they operate on a representation geometry that is too simple for robust data removal.

Lu: They’ve established that the key to better unlearning lies in explicitly targeting those minor components, which means we need new methods like MCU to ensure changes happen where they are most difficult to reverse <ref:2605.11685#pg1>.

Meng: From an engineering standpoint, the implication is that future work needs to integrate this component-aware projection directly into the unlearning loss calculation, rather than treating it as a separate add-on step.

Lalam: For the broader world, this suggests that we can move toward a more trustworthy ecosystem where data provenance and privacy are not just policies but are enforced by the very mathematical structure of how AI models forget information.

Tom: I think it means that true model erasure is going to require a deeper understanding of the internal architecture than we currently have, specifically concerning these representation subspaces.

Jane: It suggests that simply maximizing loss on a forget set isn't enough; we need to ensure the resulting parameter updates are steered away from the easily reversible directions and into those minor component spaces.

Lu: The future work they point toward involves rigorously applying this geometry analysis across various model sizes, like Gemma2-9B and Qwen3-8B, to confirm the generality of these findings <ref:2605.11685#pg1>.

Meng: I’m focused on scalability; if this approach works well for the smaller models they tested, we need to figure out how to make that projection operator efficient enough for massive foundation models.

Lalam: If we can achieve this level of robustness, it could fundamentally alter the trust relationship between users and large language models because the risk of latent knowledge reappearing is significantly reduced.

Shanghai University of Finance and Economics · Alibaba Group

cs.CL

Submitted: 2026-05-12

Updated: 2026-10-03

Code: https://github.com/sustech-nlp/MCU

Importance score: 92/100

The gist: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining, but recent studies reveal that unlearned models rapidly recover

Key concepts

Dominant Components
These are the principal directions in the model's internal representations that capture the vast majority of the total variance. Current unlearning methods primarily focus their changes on these directions, making them vulnerable to recovery attacks.
Minor Components
These represent smaller, less significant directions in the representation space that contribute less to overall variance. The study found these components are inherently more resistant to being reversed or recovered during relearning attacks.
Minor Component Unlearning (MCU)
A novel technique that modifies unlearning processes by projecting them onto the subspace of minor components. This forces the unlearning pressure into directions that are naturally harder for an attacker to reverse, significantly boosting model robustness.

Terminology

Summary

Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining, but recent studies reveal that unlearned models rapidly recover “forgotten” knowledge through relearning attacks, raising serious security concerns. This work investigates the fundamental mechanism underlying this fragility from a representation geometry perspective and proposes Minor Component Unlearning (MCU), a novel approach that explicitly targets minor components in representations to achieve substantially improved resistance to relearning attacks.

The gist

Existing unlearning methods predominantly optimize along dominant component directions, leaving minor components largely unchanged; critically, during relearning attacks, the modifications in these dominant components are easily reversed, whereas minor components exhibit stronger resistance to such reversal.

Problem Definition and Existing Vulnerabilities

LLM unlearning seeks updated parameters that ensure the mutual information between model weights and a forget set approaches zero while maintaining performance on a retain set. Current unlearning methods include Gradient Ascent (GA), Negative Preference Optimization (NPO), Representation Misdirection for Unlearning (RMU), and MLP Breaking. However, rigorous evaluations have revealed that state-of-the-art unlearning methods achieve recovery rates exceeding 88% after relearning attacks, demonstrating that they fail to truly remove knowledge from model weights. Furthermore, fine-tuning on benign, unrelated downstream tasks can inadvertently undo the unlearning effects for open-weight models.

Representation Geometry Analysis

The study investigates how unlearning and relearning affect internal representations by applying Principal Component Analysis (PCA) to MLP activations across all layers of Llama-3.1-8B on the forget set Df. This analysis yielded three key empirical observations:

  1. LLM Representations are Concentrated in Dominant Components: The first few dominant components capture the overwhelming majority of the total variance, while minor components form a long tail of small but non-negligible contributions.

  2. Unlearning Predominantly Modifies Dominant Components: Unlearning induces disproportionately large changes along the leading principal components, while minor components remain largely unchanged.

  3. Dominant Components are More Easily Recovered During Relearning: Dominant components attain substantially higher recovery ratios (often >90%) than minor components, with the same pattern across unlearning losses.

Theoretical Mechanism of Fragility

The empirical patterns are formalized through a linearized (NTK-style) analysis, leading to two theorems that explain the fragility. Theorem 1 states that after T unlearning steps, the change-ratio mirrors the explained-variance profile of Figure 2a: the unlearning update is channeled through directions that account for most of the representation variance. Theorem 2 explains recoverability: Dominant components saturate within a few attack steps, while minor components require O(1/σ 2 k) steps and remain effectively unrecovered under any bounded budget. This demonstrates that the unlearning effect is concentrated in exactly the directions the attacker can recover most cheaply.

Minor Component Unlearning (MCU) Proposal

Building on these findings, Minor Component Unlearning (MCU) is proposed as a novel approach that explicitly targets minor components of internal representations. The methodology involves:

  1. Extracting principal components from the original model’s representations on the forget set Df to identify the top-K dominant directions.

  2. Defining a projection operator, P⊥(h), to remove these top-K principal directions from any representation h, thereby confining unlearning effects to the minor component subspace.

  3. Applying this projection before computing the unlearning loss (e.g., for RMU or MLP Breaking). For example, in RMU-MCU, the loss is calculated using P⊥(h(t)) − c · u, ensuring that unlearning pressure exclusively along these robust directions.

Experimental Validation and Results

Extensive experiments on WMDP-Cyber, WMDP-Bio, and Years datasets validate MCU. Table 1 shows that MCU consistently improves robustness across base methods. When combined with CIR (Collapse of Irrelevant Representations), the results are even stronger: adding MCU further reduces the relearning gap, demonstrating that the two techniques address different aspects of the unlearning robustness problem. The analysis confirms that the change-ratio mass is concentrated in the first few PCs for all tested losses, and MCU successfully shifts this distribution toward later bins, indicating that changes are redistributed to minor components. Furthermore, experiments on Gemma2-9B and Qwen3-8B confirm the generality of these findings. The results show that MCU provides the strongest robustness, with the relearning gap ∆ being several times smaller than MLP Breaking + CIR on both datasets.

Conclusion

The research concludes that existing unlearning methods are fragile because they modify dominant components, which are easily reversed during relearning attacks. By redirecting unlearning into the minor-component subspace via MCU, the method leverages the inherent resistance of these directions to recovery, achieving substantially improved robustness while maintaining model utility.

Improvements for AI systems

Based on the scientific paper Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter, here are the specific improvements that can be made to AI systems, along with a description of what these improved systems will be able to do.


The primary improvement is the transition from current fragile unlearning methods to a representation-geometry-aware approach, specifically by implementing the proposed Minor Component Unlearning (MCU) method.

Here are the specific improvements:

  1. The AI system should implement an unlearning strategy that explicitly targets and modifies only the minor principal components of internal representations during the unlearning process, rather than optimizing along dominant components as current methods do.

  2. This requires a pre-unlearning step to extract principal components (via SVD) from the forget set representations and then projecting these vectors out before computing the unlearning loss (for RMU and MLP Breaking losses).

  3. The system must be designed to treat the mean representation as a special 0th component to be explicitly removed from this projection subspace.

  4. This modification should be applied to both representation-level unlearning objectives (RMU) and output-level unlearning objectives (MLP Breaking), ensuring that the resulting gradient updates are constrained primarily to the minor component subspace.

The improved AI system can achieve the following specific capabilities:

  1. It will exhibit significantly enhanced robustness against relearning attacks. Specifically, when an adversary attempts to recover forgotten knowledge by fine-tuning (or using other relevant datasets) on benign, unrelated downstream tasks, the recovery rate of unlearned information will be substantially lower than current state-of-the-art methods.

  2. It will maintain high utility on unrelated tasks (measured by metrics like MMLU and WikiText loss), ensuring that the knowledge suppression is precise and does not lead to catastrophic performance degradation.

  3. It will provide a more reliable guarantee of information erasure, moving beyond merely hiding knowledge to achieving a state where the mutual information between weights and the forget set approaches zero.

  4. It will be particularly effective in open-weight models, as it directly counters the vulnerability identified in these models where downstream actors can easily reverse unlearning through minimal fine-tuning.

  5. The system's performance will be superior across different model architectures (e.g., Llama-3.1-8B, Gemma2-9B, Qwen3-8B), confirming the generality of the technique across model families and not just a specific implementation detail of one architecture.

In summary, the improved system will be a significantly more secure and reliable machine learning model capable of performing targeted data suppression without leaving easily reversible traces for malicious actors.

Sources

Related papers