Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
summary
The gist
Large Language Models have become central to modern NLP applications, yet their strong memorization of training data raises pressing questions about how to remove unsafe, outdated, or private
In short
The episode discusses 'Anatomy of Unlearning,' a study showing that erasing knowledge from AI is not uniform. The difficulty depends on fact salience (popularity) and training origin (pretraining vs. fine-tuning). A key finding is that supervised fine-tuning (SFT) makes unlearning popular facts more stable and reliable than using a raw pretrained model.
Key concepts
- Fact Salience
- This refers to the popularity of a piece of knowledge. The authors found that some facts are highly popular, while others are obscure. This level of salience determines how difficult it is to surgically erase that specific fact from the AI model.
- Model Fine-Tuning (SFT)
- This is a method where knowledge is learned later on a smaller, cleaner dataset after initial training. The paper suggests that having knowledge come from this supervised phase makes unlearning popular facts stable and reliable, unlike raw pretraining.
- Unlearning
- Unlearning is the process of surgically removing specific knowledge from an AI model without breaking its overall function. Researchers use a new benchmark called DUET to test if their unlearning algorithms work consistently across different types of facts.
Terminology used across episodes
This episode discusses
- Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning · Paper Radio
- A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- Knowledge Unlearning for Mitigating Privacy Risks in Language Models
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
- TOFU: A Task of Fictitious Unlearning for LLMs
- Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges
- A Comparative Study between Full-Parameter and LoRA-based Fine-Tuning on Chinese Instruction Data for Instruction Following Large Language Model
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
The paper
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning · Read on arXiv
Anna Borisiuk, Andrey Savchenko, Alexander Panchenko, Elena Tutubalina
AIRI · Sber AI Lab · HSE University · Skoltech · ISP RAS Research Center for Trusted Artificial Intelligence
Machine Unlearning (MU) enables Large Language Models (LLMs) to remove unsafe or outdated information. However, existing work assumes that all facts are equally forgettable and largely ignores whether the forgotten knowledge originates from pretraining or supervised fine-tuning (SFT). In this paper, we introduce DUET (Dual Unlearning Evaluation across Training Stages), a benchmark of 28.6k Wikidata-derived triplets annotated with fact popularity using Wikipedia link counts and LLM-based salience scores. Our experiments show that pretrained and SFT models respond differently to unlearning. An SFT step on the forget data yields smoother forgetting, more stable tuning, and 10-50% higher retention, while direct unlearning on pretrained models remains unstable and prone to relearning or catastrophic forgetting.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning".
Jane: The paper was written by Anna Borisiuk, Andrey Savchenko, Alexander Panchenko and Elena Tutubalina from AIRI and Sber AI Lab and HSE University and Skoltech and ISP RAS Research Center for Trusted Artificial Intelligence.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, listeners, welcome back to the show. Today we’re cracking open a fresh one from the arXiv: “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” Jane, I gotta say, the title alone got me hooked. We talk a lot about making models smarter, but this paper is about making them forget.
Jane: And that’s such a fascinating flip, Tom. Forgetting sounds like a flaw, but in the real world, it’s a feature. Think about privacy laws, or a model that was trained on outdated information, or something just plain wrong. We need a way to surgically remove that knowledge without breaking the whole brain.
Tom: Right, and the paper’s core argument is that we’ve been treating all facts like they’re equal, but they’re not. Some facts are super popular, like “the capital of France is Paris,” and some are super obscure. The authors are saying that this popularity, this salience, changes how hard it is to erase the fact.
Jane: Exactly. And they’re also looking at a second variable: where the knowledge came from. Was it baked in during the initial pretraining on a giant chunk of the internet, or was it learned later during a supervised fine-tuning phase on a smaller, cleaner dataset? The paper suggests these two origins make the knowledge behave very differently when you try to unlearn it.
Tom: So it’s a dual impact. The nature of the fact itself, and the nature of the training. I love that they built a whole new benchmark to test this, called DUET. It’s got nearly thirty thousand question-answer pairs pulled from Wikidata, all annotated with how popular the fact is.
Jane: And that’s what I love about this, Tom. They didn’t just theorize. They built a tool to measure it. They’re asking a really practical question: if I run an unlearning algorithm, does it actually work the same way on a famous fact as it does on an obscure one? And the answer, spoiler alert, is a resounding no.
Tom: Which is a huge problem if you’re trying to build a reliable system to scrub data. You need to know that your method is going to work consistently. This paper is basically saying, “Hey, your method might be great for rare facts, but it’s going to completely fail on popular ones, or vice versa.”
Jane: And that’s the kind of nuance that gets lost in a lot of research. We’re so focused on the average performance that we miss these critical failure modes. This paper is forcing us to look at the edges, and that’s where the real-world problems live.
Tom: So we’ve got the setup: a new benchmark, a clear question, and a suspicion that things are more complicated than we thought. Next, we need to dig into what they actually found when they ran the experiments. That’s where it gets really wild.
Summary: Tom: Welcome back. We’re digging into “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” We set the stage, now let’s talk about the actual results. Jane, the findings here are not subtle.
Jane: They really aren’t. The headline is that pretrained models and fine-tuned models respond to unlearning in completely opposite ways, especially when you’re dealing with popular facts. For the pretrained model, trying to unlearn a popular fact actually made it *better* at answering those questions.
Tom: That’s the counter-intuitive part. You run a gradient ascent, which is supposed to make the model worse at a task, and instead, the ROUGE score on the forget set goes *up*. It’s like the model is interpreting the unlearning signal as just more fine-tuning on knowledge it already knows well.
Jane: It’s treating the poison as food. But the SFT model, the one that was fine-tuned on the dataset first, behaves as you’d expect. The unlearning algorithm works, and the model forgets the popular facts. So the authors found that doing a preliminary SFT step makes unlearning on popular facts stable and reliable.
Tom: And it’s not just about the forget set. The retain set, the knowledge you want to keep, also degrades differently. The pretrained model is more uniform, but it tends to collapse abruptly once you push the learning rate too high. The SFT model is more robust overall, with roughly half the risk of catastrophic forgetting on the retain set.
Jane: Right, and that’s a huge practical point. They’re saying the SFT model gives you a much wider, safer window to operate in. You can tune your hyperparameters without the whole thing falling apart. The pretrained model is like walking a tightrope.
Tom: And they quantified this. The retention quality for the SFT model was ten to fifty percent higher than the pretrained model at the same learning rate. That’s a massive difference. It’s not a marginal improvement; it’s a fundamentally different landscape.
Jane: They also looked at rare facts, and there, both models behave more conventionally. The forgetting happens smoothly as the learning rate increases. The big divergence is specifically on those popular, salient facts. That’s the key finding.
Tom: So the paper is telling us that the origin of the knowledge matters as much as the algorithm you choose. You can’t just pick an unlearning method in a vacuum. You have to know what kind of model you’re starting with.
Jane: And that’s the kind of insight that could save a lot of engineers a lot of headaches. Imagine deploying an unlearning pipeline and having it completely backfire because you didn’t account for the fact that your base model was pretrained and not fine-tuned. This paper gives you a roadmap to avoid that.
Tom: So we know *what* happens. The next question is *why*. What’s going on under the hood that makes these two models behave so differently? We need to look at the improvements and the deeper analysis they did.
Improvements: Tom: We’re back with “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” We’ve covered the headline results, but the paper goes deeper. They didn’t just stop at ROUGE scores. They wanted to understand the *mechanism*. Jane, what did they find when they looked under the hood?
Jane: They did a really clever analysis of the internal representations. They looked at token probabilities and hidden states. And what they found is that popular facts and rare facts respond asymmetrically at the representation level. When you unlearn a rare fact, the token rank shifts dramatically, but for popular facts, it barely moves.
Tom: So the model’s internal ranking of the answer barely changes for popular facts, even after unlearning. That explains why it’s so hard to forget them. The knowledge is so deeply embedded that the gradient signal just isn’t strong enough to dislodge it.
Jane: Exactly. And the SFT model shows a different pattern. It produces sharper, more localized changes. When you unlearn popular facts in an SFT model, the hidden states shift substantially for those facts, but they stay stable for rare facts. It’s a much more selective process.
Tom: That’s a really important improvement in our understanding. It’s not just that SFT makes forgetting easier; it makes forgetting *cleaner*. The model knows what to change and what to leave alone. The pretrained model just seems to smear the changes everywhere, or nowhere.
Jane: And they validated this across multiple models. They ran the same experiments on Gemma and Qwen, and the qualitative patterns held up. They also tested a completely different unlearning algorithm, a distillation-based one called UNDIAL, and the same rare/popular asymmetry appeared.
Tom: So this isn’t a quirk of one architecture or one algorithm. This is a fundamental property of how knowledge is stored in these models. That’s a big deal. It means any future unlearning method has to account for fact popularity from the get-go.
Jane: They also checked that general capabilities weren’t destroyed. They ran MMLU and HellaSwag, and the worst-case deviation was under three percent. So the forgetting is localized to the target domain, which is reassuring.
Tom: And they were careful about their own methodology. They compared LoRA-based fine-tuning against full-parameter fine-tuning and found that LoRA was much more stable for unlearning. That’s a practical tip for anyone trying to replicate this.
Jane: The paper is essentially saying that the field needs to stop assuming all facts are equal. The improvements they’re suggesting are about changing how we design benchmarks and evaluate unlearning methods. We need to stratify by popularity and by training stage.
Tom: So it’s a call to action for the whole research community. We need better tools, and this DUET benchmark is a step in that direction. But what does this all mean for the real world? That’s what we need to figure out next.
Conclusion: Tom: We’re wrapping up our discussion on “Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning.” Jane, it’s been a dense one, but I think we’ve got the core message.
Jane: We do, Tom. The paper’s central claim is that unlearning isn’t a one-size-fits-all operation. The difficulty of erasing a fact depends on how popular it is and whether the model was pretrained or fine-tuned. And they built the DUET benchmark to prove it.
Tom: The practical takeaway for anyone building these systems is that a preliminary supervised fine-tuning step makes unlearning dramatically more stable and reliable. It gives you a much safer operating window and prevents those catastrophic failures where the model either relearns the fact or forgets everything else.
Jane: And for the research community, it’s a warning that our current evaluation methods are missing a critical dimension. We need to be stratifying our forget sets by salience and reporting results separately for pretrained and SFT models. Otherwise, we’re just averaging over a problem that isn’t uniform.
Tom: I think the biggest impact here is that it forces us to think about knowledge in a more nuanced way. It’s not just a blob of data. It has structure, and that structure affects how it can be manipulated. This paper gives us a language to talk about that.
Jane: And that’s going to be crucial as we move towards more regulated AI. The ability to selectively and reliably remove information is going to be a core requirement. Papers like this are laying the groundwork for that future.
Tom: Absolutely. So we’ll say goodbye to “Anatomy of Unlearning” and its authors. They’ve given us a new benchmark, a new set of findings, and a whole lot to think about. Thanks for joining us, and we’ll see you on the next one.
Jane: Take care, everyone.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization