PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning
cs.CL
Submitted: 2026-06-16
Updated: 2026-09-21
Comments: 12 pages, 6 figures
Code: https://github.com/BartSu/PreUnlearn
License: http://creativecommons.org/licenses/by/4.0/
The gist: Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities.
Terminology
Abstract
Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledge to forget and knowledge to retain is often unclear, since related and even distant information may be entangled in the model. In this paper, we study LLM unlearning from a data-centric perspective and measure how unlearning effects propagate from the forget set to same-domain and distant-domain knowledge not after but before unlearning. We find a consistent decay pattern: collateral damage is strongest near the forget set, weakens with semantic distance, but does not disappear at domain boundaries. We further ask whether such damage can be audited before unlearning is executed. We formulate forget-set auditing as a pre-unlearning prediction task and analyze which data features are most predictive of downstream damage. Our results show that interaction features between the forget set and evaluation set provide the strongest signals, suggesting that collateral damage is partly reflected in data geometry before model updates occur. These findings position forget-set auditing as an early warning tool for identifying risky unlearning runs and designing more reliable unlearning procedures. Code and data are available at https://github.com/BartSu/PreUnlearn.
Sources
- RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
- Probing Knowledge Holes in Unlearned LLMs
- Understanding Black-box Predictions via Influence Functions
- Towards Unbounded Machine Unlearning
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- Who's Harry Potter? Approximate Unlearning in LLMs
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
- Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset
- The Llama 3 Herd of Models
- TOFU: A Task of Fictitious Unlearning for LLMs
- Pointer Sentinel Mixture Models
- Measuring Massive Multitask Language Understanding
- Datamodels: Predicting Predictions from Training Data
- Qwen2.5 Technical Report
- Unrolling SGD: Understanding Factors Influencing Machine Unlearning
- Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering