Can an AI Assistant Really Forget? Auditable Deletion from Addressable Memory
cs.LG, cs.CL
Submitted: 2026-07-30
Updated: 2026-09-11
Comments: 19 pages. Major revision focusing on frozen-Gemma addressable memory, deletion certificates, and behavioral audits, including a decoded LongMemEval follow-up. Native KDA transport and replay are presented in a separate companion paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Certifying that a deletion did what it declared does not certify that the record left no trace: a small distance to the implementation's own reference does not imply a small distance to the state
Terminology
Abstract
Certifying that a deletion did what it declared does not certify that the record left no trace: a small distance to the implementation's own reference does not imply a small distance to the state that never stored the record. This paper installs a deletion interface into a pretrained language model and measures both distances. We retrofit a support-vector memory gate into the global attention layers of a frozen Gemma 3 without changing a weight. Each stored record owns a set of rows, and deleting it removes those rows and re-solves only the storage problems they touched. At 4B the retrofit admits exactly the records the base model recalls, at a paired perplexity cost under 2%; the same recipe fails at 1B and 12B, which we report. Every executed deletion agreed with an independently reconstructed reference on every registered probe, and under sampling, targeted elicitation, related-data relearning, and membership inference an edited record was about as hard to extract as one never stored, while a prompt instruction to ignore the same record left it fully extractable. On 96 long conversational histories with decoded answers, the edited assistant disclosed the deleted record in 15 histories against 13 for a rebuild that never stored it and 54 for the instruction, and a blinded review of the outputs the matcher had cleared found that its misses were aliases or normalization failures of the answer, with no paraphrase among them. The edit also suppressed the deleted answer below the never-stored level, a signature an auditor can read. The result is a retrofit that makes a frozen model's memory addressable, a certificate for what the retrofit does, and a measurement of the distance that remains to the stronger guarantee; which of the two a system can offer is decided when the memory is written.
Sources
- Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
- Machine Unlearning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Gemma 3 Technical Report
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- LoRA: Low-Rank Adaptation of Large Language Models
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention
- SnapKV: LLM Knows What You are Looking for Before Generation
- Eight Methods to Evaluate Robust Unlearning in LLMs
- TOFU: A Task of Fictitious Unlearning for LLMs
- Pointer Sentinel Mixture Models
- In-Context Unlearning: Language Models as Few Shot Unlearners
- What a Deletion Certificate Covers, and Where It Expires: Auditable Removal from a Support-Vector Memory
- Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks