What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs
cs.CL, cs.AI
Submitted: 2025-10-29
Updated: 2026-08-26
Comments: 17 pages, 6 figures, 12 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Unlearning in large language models (LLMs) is usually evaluated as whether an "unlearned" fact can be recovered.
Terminology
Abstract
Unlearning in large language models (LLMs) is usually evaluated as whether an "unlearned" fact can be recovered. We instead ask whether a fact's structural entanglement with the rest of a model's knowledge predicts whether it leaks after unlearning, whether this relationship changes systematically, and whether it is causal. Across varied unlearning algorithms (WHP and GA+KL), two domains (both fictional Harry Potter and non-fictional U.S. Senators, 2000-2010), and four models (2.7B-13B parameters), we find that before unlearning, more entangled facts are recalled more often (r = +0.39 to +0.51). WHP weakens this relationship but stays positive (r = +0.16 to +0.33); GA+KL inverts it in every domain and model size (r = -0.14 to-0.25). To our knowledge, this is the first report of this specific reversal in the unlearning literature. To confirm this is causal beyond correlation, we directly manipulate a prompt's entanglement score, holding its content, target model fixed, and show recall moves in the predicted direction and reverses sign under GA+KL in the same direction as the correlational analysis. This manipulation is our central evidence that unlearning acts on the underlying knowledge structure, not just its output. Building on this, we train a predictive model that estimates a prompt's post-unlearning factuality/hallucination profile before unlearning is run, creating a unique way of model auditing in triaging which prompts are likely to leak.
Sources
- TOFU: A Task of Fictitious Unlearning for LLMs
- Extracting Training Data from Large Language Models
- Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4
- Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
- Who's Harry Potter? Approximate Unlearning in LLMs
- Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Copyright Violations and Large Language Models
- Machine Unlearning: A Survey
- The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- RUB: Evaluating Residual Knowledge in Unlearned Models
- Disentangling Knowledge Representations for Large Language Model Editing
- OPT: Open Pre-trained Transformer Language Models
- Attention Heads of Large Language Models: A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering