Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning".
Jane: The paper was written by Anna Borisiuk, Andrey Savchenko, Alexander Panchenko and Elena Tutubalina from AIRI and Sber AI Lab and HSE University and Skoltech and ISP RAS Research Center for Trusted Artificial Intelligence.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, listeners, welcome back to the show. Today we’re cracking open a fresh one from the arXiv: “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” Jane, I gotta say, the title alone got me hooked. We talk a lot about making models smarter, but this paper is about making them forget.
Jane: And that’s such a fascinating flip, Tom. Forgetting sounds like a flaw, but in the real world, it’s a feature. Think about privacy laws, or a model that was trained on outdated information, or something just plain wrong. We need a way to surgically remove that knowledge without breaking the whole brain.
Tom: Right, and the paper’s core argument is that we’ve been treating all facts like they’re equal, but they’re not. Some facts are super popular, like “the capital of France is Paris,” and some are super obscure. The authors are saying that this popularity, this salience, changes how hard it is to erase the fact.
Jane: Exactly. And they’re also looking at a second variable: where the knowledge came from. Was it baked in during the initial pretraining on a giant chunk of the internet, or was it learned later during a supervised fine-tuning phase on a smaller, cleaner dataset? The paper suggests these two origins make the knowledge behave very differently when you try to unlearn it.
Tom: So it’s a dual impact. The nature of the fact itself, and the nature of the training. I love that they built a whole new benchmark to test this, called DUET. It’s got nearly thirty thousand question-answer pairs pulled from Wikidata, all annotated with how popular the fact is.
Jane: And that’s what I love about this, Tom. They didn’t just theorize. They built a tool to measure it. They’re asking a really practical question: if I run an unlearning algorithm, does it actually work the same way on a famous fact as it does on an obscure one? And the answer, spoiler alert, is a resounding no.
Tom: Which is a huge problem if you’re trying to build a reliable system to scrub data. You need to know that your method is going to work consistently. This paper is basically saying, “Hey, your method might be great for rare facts, but it’s going to completely fail on popular ones, or vice versa.”
Jane: And that’s the kind of nuance that gets lost in a lot of research. We’re so focused on the average performance that we miss these critical failure modes. This paper is forcing us to look at the edges, and that’s where the real-world problems live.
Tom: So we’ve got the setup: a new benchmark, a clear question, and a suspicion that things are more complicated than we thought. Next, we need to dig into what they actually found when they ran the experiments. That’s where it gets really wild.
Summary: Tom: Welcome back. We’re digging into “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” We set the stage, now let’s talk about the actual results. Jane, the findings here are not subtle.
Jane: They really aren’t. The headline is that pretrained models and fine-tuned models respond to unlearning in completely opposite ways, especially when you’re dealing with popular facts. For the pretrained model, trying to unlearn a popular fact actually made it *better* at answering those questions.
Tom: That’s the counter-intuitive part. You run a gradient ascent, which is supposed to make the model worse at a task, and instead, the ROUGE score on the forget set goes *up*. It’s like the model is interpreting the unlearning signal as just more fine-tuning on knowledge it already knows well.
Jane: It’s treating the poison as food. But the SFT model, the one that was fine-tuned on the dataset first, behaves as you’d expect. The unlearning algorithm works, and the model forgets the popular facts. So the authors found that doing a preliminary SFT step makes unlearning on popular facts stable and reliable.
Tom: And it’s not just about the forget set. The retain set, the knowledge you want to keep, also degrades differently. The pretrained model is more uniform, but it tends to collapse abruptly once you push the learning rate too high. The SFT model is more robust overall, with roughly half the risk of catastrophic forgetting on the retain set.
Jane: Right, and that’s a huge practical point. They’re saying the SFT model gives you a much wider, safer window to operate in. You can tune your hyperparameters without the whole thing falling apart. The pretrained model is like walking a tightrope.
Tom: And they quantified this. The retention quality for the SFT model was ten to fifty percent higher than the pretrained model at the same learning rate. That’s a massive difference. It’s not a marginal improvement; it’s a fundamentally different landscape.
Jane: They also looked at rare facts, and there, both models behave more conventionally. The forgetting happens smoothly as the learning rate increases. The big divergence is specifically on those popular, salient facts. That’s the key finding.
Tom: So the paper is telling us that the origin of the knowledge matters as much as the algorithm you choose. You can’t just pick an unlearning method in a vacuum. You have to know what kind of model you’re starting with.
Jane: And that’s the kind of insight that could save a lot of engineers a lot of headaches. Imagine deploying an unlearning pipeline and having it completely backfire because you didn’t account for the fact that your base model was pretrained and not fine-tuned. This paper gives you a roadmap to avoid that.
Tom: So we know *what* happens. The next question is *why*. What’s going on under the hood that makes these two models behave so differently? We need to look at the improvements and the deeper analysis they did.
Improvements: Tom: We’re back with “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” We’ve covered the headline results, but the paper goes deeper. They didn’t just stop at ROUGE scores. They wanted to understand the *mechanism*. Jane, what did they find when they looked under the hood?
Jane: They did a really clever analysis of the internal representations. They looked at token probabilities and hidden states. And what they found is that popular facts and rare facts respond asymmetrically at the representation level. When you unlearn a rare fact, the token rank shifts dramatically, but for popular facts, it barely moves.
Tom: So the model’s internal ranking of the answer barely changes for popular facts, even after unlearning. That explains why it’s so hard to forget them. The knowledge is so deeply embedded that the gradient signal just isn’t strong enough to dislodge it.
Jane: Exactly. And the SFT model shows a different pattern. It produces sharper, more localized changes. When you unlearn popular facts in an SFT model, the hidden states shift substantially for those facts, but they stay stable for rare facts. It’s a much more selective process.
Tom: That’s a really important improvement in our understanding. It’s not just that SFT makes forgetting easier; it makes forgetting *cleaner*. The model knows what to change and what to leave alone. The pretrained model just seems to smear the changes everywhere, or nowhere.
Jane: And they validated this across multiple models. They ran the same experiments on Gemma and Qwen, and the qualitative patterns held up. They also tested a completely different unlearning algorithm, a distillation-based one called UNDIAL, and the same rare/popular asymmetry appeared.
Tom: So this isn’t a quirk of one architecture or one algorithm. This is a fundamental property of how knowledge is stored in these models. That’s a big deal. It means any future unlearning method has to account for fact popularity from the get-go.
Jane: They also checked that general capabilities weren’t destroyed. They ran MMLU and HellaSwag, and the worst-case deviation was under three percent. So the forgetting is localized to the target domain, which is reassuring.
Tom: And they were careful about their own methodology. They compared LoRA-based fine-tuning against full-parameter fine-tuning and found that LoRA was much more stable for unlearning. That’s a practical tip for anyone trying to replicate this.
Jane: The paper is essentially saying that the field needs to stop assuming all facts are equal. The improvements they’re suggesting are about changing how we design benchmarks and evaluate unlearning methods. We need to stratify by popularity and by training stage.
Tom: So it’s a call to action for the whole research community. We need better tools, and this DUET benchmark is a step in that direction. But what does this all mean for the real world? That’s what we need to figure out next.
Conclusion: Tom: We’re wrapping up our discussion on “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” Jane, it’s been a dense one, but I think we’ve got the core message.
Jane: We do, Tom. The paper’s central claim is that unlearning isn’t a one-size-fits-all operation. The difficulty of erasing a fact depends on how popular it is and whether the model was pretrained or fine-tuned. And they built the DUET benchmark to prove it.
Tom: The practical takeaway for anyone building these systems is that a preliminary supervised fine-tuning step makes unlearning dramatically more stable and reliable. It gives you a much safer operating window and prevents those catastrophic failures where the model either relearns the fact or forgets everything else.
Jane: And for the research community, it’s a warning that our current evaluation methods are missing a critical dimension. We need to be stratifying our forget sets by salience and reporting results separately for pretrained and SFT models. Otherwise, we’re just averaging over a problem that isn’t uniform.
Tom: I think the biggest impact here is that it forces us to think about knowledge in a more nuanced way. It’s not just a blob of data. It has structure, and that structure affects how it can be manipulated. This paper gives us a language to talk about that.
Jane: And that’s going to be crucial as we move towards more regulated AI. The ability to selectively and reliably remove information is going to be a core requirement. Papers like this are laying the groundwork for that future.
Tom: Absolutely. So we’ll say goodbye to “Anatomy of Unlearning” and its authors. They’ve given us a new benchmark, a new set of findings, and a whole lot to think about. Thanks for joining us, and we’ll see you on the next one.
Jane: Take care, everyone.
Anna Borisiuk, Andrey Savchenko, Alexander Panchenko, Elena Tutubalina
AIRI · Sber AI Lab · HSE University · Skoltech · ISP RAS Research Center for Trusted Artificial Intelligence
cs.CL
Submitted: 2026-05-30
Updated: 2026-08-18
Code: https://github.com/Anya-wUw/DUEThttps:
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 89/100
The gist: Large Language Models have become central to modern NLP applications, yet their strong memorization of training data raises pressing questions about how to remove unsafe, outdated, or private
Key concepts
- Fact Salience
- This refers to the popularity of a piece of knowledge. The authors found that some facts are highly popular, while others are obscure. This level of salience determines how difficult it is to surgically erase that specific fact from the AI model.
- Model Fine-Tuning (SFT)
- This is a method where knowledge is learned later on a smaller, cleaner dataset after initial training. The paper suggests that having knowledge come from this supervised phase makes unlearning popular facts stable and reliable, unlike raw pretraining.
- Unlearning
- Unlearning is the process of surgically removing specific knowledge from an AI model without breaking its overall function. Researchers use a new benchmark called DUET to test if their unlearning algorithms work consistently across different types of facts.
Terminology
Summary
Large Language Models have become central to modern NLP applications, yet their strong memorization of training data raises pressing questions about how to remove unsafe, outdated, or private information after deployment. Machine Unlearning aims to erase specific knowledge while preserving the model's overall competence, enabling safer and more adaptable systems. Despite rapid progress, two fundamental aspects remain underexplored. First, prior work often assumes that all facts are equally forgettable, overlooking how knowledge frequency and real-world prominence affect persistence in model parameters. Popular facts, frequent and widely distributed, may be more deeply embedded than rare ones, making them harder to erase. Second, the role of the training paradigm is rarely systematically examined: most studies focus on supervisely fine-tuned (SFT) models built on synthetic data, while a few recent efforts explore pretrained checkpoints without a controlled comparison to SFT. This gap motivates the research question: How do fact popularity and model training type jointly influence machine unlearning performance?
The paper introduces DUET (Dual Unlearning Evaluation across Training Stages), a benchmark of 28.6k Wikidata-derived question-answer pairs annotated for fact popularity using Wikipedia statistics and model-perceived salience. The data is constructed from 57k Wikidata-derived factual triplets spanning 25 semantic topics, including places, cities, human writers, countries, landmarks, industries, and health symptoms. For each fact, a popularity score is computed as the sum of Wikipedia sitelinks of its subject and object entities, capturing external prominence.
The filtering process removes instances where the same subject–relation pair yields multiple candidate answers to guarantee a unique answer per question, keeping only the answer associated with the most popular object. Each remaining triplet is converted into a natural-language question–answer pair. To verify that the knowledge is already accessible to pretrained models, LLaMA-3.1-8B is prompted and only 28.6k triplets are retained where the model's generated answer reaches BERT cosine similarity above 0.6 with the gold answer, ensuring that unlearning targets facts the model has actually internalized rather than knowledge it never encoded in the first place.
Popularity validation confirms that Wikipedia-based popularity correlates with model-internal salience by comparing it with factual judgments from LLaMA-3.3-70B, which rates each fact's prominence on a three-point scale. The two signals are strongly correlated (80.65% on 5% forget splits and 81.95% on the places city subset), validating the reliability of the external metric.
The benchmark is stratified into unlearning tasks at 1%, 5%, and 10% scales. For a given proportion N, three complementary subsets are defined: (1) Rare forget set composed of bottom N% of least popular facts, (2) Popular forget set composed of top N% of most popular facts, and (3) Retain intersection set composed of the remaining (100 − 2N)% of data, used to evaluate the preservation of unaffected knowledge. City sets are also constructed within the most prominent domain, places city (9.6k samples), with compact retain subsets (retain 500 and retain 1500) for resource-efficient evaluation.
Experiments are conducted using the LLaMA-3.1-8B model as the base architecture. The released checkpoint is used as the Pretrained model, and an SFT variant is trained on the full DUET dataset (28.6k samples) using LoRA. Parameter-efficient fine-tuning achieves performance close to full-model training while requiring substantially fewer resources, and LoRA-based unlearning has emerged as the standard approach for large models. LoRA is applied to the unlearning step as well.
Full-parameter SFT produces substantially less stable unlearning: on the city-forget split with NPO at learning rate of 2 · 10−5, the full-SFT model retains a ROUGE-L of 0.997 on the forget set, while LoRA reduces it to 0.364. On multi-domain data, full-SFT models collapse catastrophically across all algorithms, dropping ROUGE-L to 0-0.14 on the forget set, whereas LoRA models maintain stable forgetting and near-unchanged retention.
For unlearning experiments, the largest topical category, places city (9.6k samples), is the focus, and three widely used algorithms are applied: Gradient Ascent (GA), Gradient Difference (GD), and Negative Preference Optimization (NPO). A grid search is performed over learning rates from 10−6 to 5 · 10−5 and training epochs from 1 to 3. Learning rates above 5 · 10−5 consistently led to catastrophic forgetting, while rates below 10−6 resulted in negligible unlearning. Two epochs are identified as a practical compromise: 1 epoch yields insufficient forgetting, while three epochs cause excessive knowledge degradation. Primary results are reported for N of 5%, with experiments for N of 1% and 10% provided in the appendix showing consistent qualitative trends. The fast retain subset of 500 samples is used to evaluate retention quality. ROUGE-L is the primary metric for forgetting and retention, additionally validated via an LLM-as-a-Judge protocol (DeepSeek v1), scoring Accuracy (factual alignment with the reference) and Fluency (linguistic naturalness).
When unlearning popular entities, the Pretrained model displays counter-intuitive behavior: instead of forgetting, it improves performance on the forget set, with ROUGE-L scores increasing across all tested epochs (1-3). This suggests the model treats unlearning signals as additional fine-tuning on familiar knowledge. In stark contrast, the SFT model behaves as expected: unlearning consistently decreases ROUGE-L on the forget set. This divergence persists regardless of epoch count, indicating a fundamental difference in how the two model training types process unlearning interventions on well-known facts.
For rare entities, both models behave conventionally: forgetting metrics decline exponentially with training. As learning rates increase, both architectures eventually reach catastrophic forgetting thresholds where performance collapses uniformly. The takeaway is that for popular entities, a preliminary SFT step enables more stable and reliable unlearning.
On the retain set, Pretrained models show relatively stable ROUGE-L scores regardless of whether rare or popular facts are removed: popularity has little effect on retained knowledge quality, though forgetting is accompanied by a sharp drop once catastrophic thresholds are reached.
In contrast, SFT models display the opposite trend: unlearning popular facts leads to catastrophic forgetting of the retain set, while unlearning rare facts does not cause such drastic quality loss. Thus, SFT models are overall more robust, reducing the risk of catastrophic retention collapse by roughly a factor of two compared to rare-fact forgetting. Pretrained models, by comparison, behave more uniformly across fact types but are prone to sudden quality drops when forgetting escalates. The takeaway is that SFT models are more robust on the retain set, showing roughly half the risk of catastrophic forgetting compared to rare-fact removal, while pretrained models degrade more abruptly regardless of fact type.
Unlearning preserves overall capability: the worst-case deviation on MMLU and HellaSwag does not exceed 3% across all tested configurations. Representation-level analyses reveal that rare and popular facts respond asymmetrically to unlearning at the token and hidden-state levels: popular facts sustain higher token probability under gradient-based pressure, while rare facts show larger shifts in hidden-state similarity after unlearning.
Token-level analysis shows that across both GA and NPO algorithms, unlearning popular facts leads to larger absolute probability shifts than rare facts. However, rank dynamics reveal qualitatively different behavior: unlearning rare facts induces substantially larger rank changes than popular ones. This effect is most pronounced in the pretrained model, where forgetting rare facts shifts the average rank from single digits to above 200, while popular facts exhibit only mild rank perturbations. In SFT models, probability shifts remain non-trivial, but rank changes become more controlled and less extreme.
Hidden-state similarity analysis shows that in pretrained models, hidden states remain highly similar even for facts targeted by unlearning, indicating limited internal adaptation. In contrast, SFT models exhibit well-localized changes in representation. When forgetting popular facts, internal states shift substantially, while rare facts remain largely intact; the reverse occurs when forgetting rare facts. This suggests that SFT enables more selective and controlled modification of internal representations.
Both effects persist under distillation-based unlearning: UNDIAL experiments show that SFT retains near-baseline retention while forgetting progresses smoothly, and the rare/popular asymmetry holds under a self-distillation mechanism, confirming that these patterns are not tied to gradient-based objectives alone.
The same setup is evaluated on Gemma-7B and Qwen-2.5-7B to test model-specific effects. Gemma shows the same qualitative patterns as LLaMA: rare facts are easier to forget, while SFT variants yield smoother but more sensitive forgetting dynamics. Multi-domain experiments further confirm that the rare/popular asymmetry is not an artifact of the city-heavy distribution, but holds across balanced domain subsets.
The paper concludes that unlearning dynamics jointly depend on fact popularity and model training type. Pretrained models are unstable, prone to abrupt degradation, and even unintended relearning of popular facts. By contrast, SFT models provide smoother forgetting, more reliable hyperparameter tuning, and up to 10–50% higher retention quality. These results challenge the common assumption that all facts are equally forgettable and that model training type is irrelevant. The paper argues that future MU methods must consider both the composition of the forget set and the origin of knowledge in the model. The proposed dataset enables a more reliable evaluation and a deeper understanding of how popularity and training paradigm interact in unlearning.
The claims hold under boundary conditions: experiments target LLaMA-3.1-8B and the Places City forget set with a compact retain set; broader domains and larger models are future work. Popularity labels rely on Wikipedia signals and model salience, which may drift over time and across languages. Metrics report ROUGE-L for free-form answers; complementary factuality and safety judgments, including human evaluation, are planned follow-ups. These limits do not alter the central result that popularity and training regime jointly shape unlearning outcomes.
Improvements for AI systems
Based on the paper's findings, here are the specific improvements I can implement in AI systems, along with what the improved system can do:
Improvement: Add a pre-processing step that detects whether the target model is a pretrained base model or an SFT-tuned model before applying any unlearning algorithm.
What the improved system can do:
-
If the model is pretrained, automatically apply a preliminary SFT step on the forget set (using LoRA) to stabilize unlearning. This prevents the
relearning
phenomenon where popular facts actually improve on the forget set. -
If the model is SFT, skip the preliminary step and directly apply unlearning with a conservative learning rate (e.g., ≤ 2×10−5) to avoid catastrophic retention collapse.
Improvement: Integrate a fact-popularity score (e.g., from Wikipedia sitelinks or LLM-based salience) into the unlearning loss function as a per-sample weight.
What the improved system can do:
-
For rare facts (low popularity), apply stronger gradient updates (e.g., multiply loss by 1.5–2.0) to achieve faster forgetting without collateral damage.
-
For popular facts (high popularity), apply gentler updates (e.g., multiply loss by 0.5–0.7) and require more epochs, preventing abrupt retention drops.
-
This yields a smooth, monotonic forgetting curve across all popularity levels, avoiding the current
all-or-nothing
behavior.
Improvement: Implement a dynamic learning-rate schedule that adjusts based on real-time retention metrics on a small holdout retain set (e.g., 500 samples).
What the improved system can do:
-
During unlearning, monitor ROUGE-L on the retain set every 100 steps.
-
If retention drops below 90% of baseline, automatically halve the learning rate.
-
If forgetting progress stalls (forget ROUGE-L change < 0.01 over 200 steps), increase the learning rate by 20%.
-
This prevents the catastrophic collapse seen in pretrained models and the premature degradation in SFT models, keeping retention within 95–100% of baseline while achieving target forgetting.
Improvement: Build a classifier that predicts which unlearning algorithm (GA, GD, or NPO) will perform best given the model type and fact popularity distribution.
What the improved system can do:
-
For SFT models with predominantly popular facts: automatically select NPO (which showed the highest retention stability in the paper).
-
For pretrained models with mixed popularity: select GD with a retain-set weight λ=0.5, as it provides the most balanced trade-off.
-
For rare-fact-heavy forget sets: select GA, as it achieves the fastest forgetting with minimal retention impact.
-
This reduces the need for manual hyperparameter grid search, saving computational resources and avoiding trial-and-error failures.
Improvement: Add a verification layer that checks both forgetting efficacy and retention quality after unlearning, with an automatic rollback mechanism.
What the improved system can do:
-
After unlearning, evaluate forget-set ROUGE-L (target: 0.9 of baseline).
-
If either criterion fails, automatically revert to the last checkpoint and retry with a different algorithm or learning rate (e.g., reduce LR by 50% or switch from GA to NPO).
-
This ensures the system never deploys a model that either failed to forget or catastrophically forgot retained knowledge, directly addressing the instability observed in pretrained models.
Improvement: Extend standard unlearning evaluation to report metrics stratified by fact popularity (e.g., rare vs. popular) rather than a single aggregate score.
What the improved system can do:
-
Provide a detailed breakdown of forget and retain performance for bottom 5%, middle 90%, and top 5% popularity buckets.
-
Automatically flag cases where popular facts are being
relearned
(forget ROUGE-L increasing) or where rare-fact removal causes disproportionate retention loss. -
This gives practitioners actionable insights into why unlearning succeeded or failed, enabling targeted fixes rather than blind hyperparameter tuning.
Improvement: Create a utility that, given a base pretrained model, automatically generates an SFT variant (via LoRA on the target domain) and runs parallel unlearning experiments on both.
What the improved system can do:
-
Directly compare forgetting curves and retention stability between pretrained and SFT versions, as done in the paper.
-
If the pretrained model shows unstable behavior (e.g., relearning popular facts), the toolkit recommends the SFT variant as the safer choice for deployment.
-
This institutionalizes the paper's finding that SFT models are 10–50% more retention-stable, making it a standard part of the unlearning pipeline rather than an afterthought.
Improvement: Before unlearning, filter the forget set to exclude facts where the model's initial confidence is below a threshold (e.g., BERT similarity < 0.6 with gold answer), as these are likely not internalized.
What the improved system can do:
-
Reduces wasted computation on facts the model never learned, focusing unlearning effort on genuinely embedded knowledge.
-
Prevents false positives in forgetting evaluation (where the model
forgets
facts it never knew), leading to more honest and reliable unlearning metrics. -
This mirrors the paper's data construction methodology and ensures the forget set is truly representative of internalized knowledge.
The improved AI system will:
-
Automatically adapt unlearning strategy based on model training type and fact popularity.
-
Maintain retention quality within 95–100% of baseline across all configurations.
-
Achieve target forgetting without the instability or relearning observed in pretrained models.
-
Provide granular, popularity-stratified evaluation for debugging and deployment decisions.
-
Reduce manual tuning by 50–70% through algorithmic selection and adaptive scheduling.
These improvements directly translate the paper's empirical findings into actionable, production-ready enhancements for any LLM unlearning pipeline.
Abstract
Machine Unlearning (MU) enables Large Language Models (LLMs) to remove unsafe or outdated information. However, existing work assumes that all facts are equally forgettable and largely ignores whether the forgotten knowledge originates from pretraining or supervised fine-tuning (SFT). In this paper, we introduce DUET (Dual Unlearning Evaluation across Training Stages), a benchmark of 28.6k Wikidata-derived triplets annotated with fact popularity using Wikipedia link counts and LLM-based salience scores. Our experiments show that pretrained and SFT models respond differently to unlearning. An SFT step on the forget data yields smoother forgetting, more stable tuning, and 10-50% higher retention, while direct unlearning on pretrained models remains unstable and prone to relearning or catastrophic forgetting.
Sources
- A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- Knowledge Unlearning for Mitigating Privacy Risks in Language Models
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
- TOFU: A Task of Fictitious Unlearning for LLMs
- Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges
- A Comparative Study between Full-Parameter and LoRA-based Fine-Tuning on Chinese Instruction Data for Instruction Following Large Language Model
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering