Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance
cs.LG, cs.AI, cs.CR, cs.NE
Submitted: 2026-05-03
Updated: 2026-09-21
Code: https://github.com/Rupawheatly/MLDU2
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent attacks show that behavioural unlearning of large language models leaves internal traces recoverable by adversarial probes.
Terminology
Abstract
Recent attacks show that behavioural unlearning of large language models leaves internal traces recoverable by adversarial probes. We characterise where this retention lives and show it can be surgically removed without measurable capability cost. Our central protocol is a leave-one-out cross-sequence probe that tests whether a memorisation signature generalises across held-out sequences. The signature is real and consistent across scale: memorisation-specific gaps of +0.32, +0.19, +0.30 on Pythia-70M, GPT-2 medium, and Mistral-7B; on Pythia-70M, the random-initialisation control collapses to-0.04 at the deepest layer where the pretrained signature peaks. The probe direction is causally separable from recall -- projecting it out collapses the signature locally (+0.44 -> -0.19) while behavioural recall barely changes -- and a probe trained on naturally memorised content does not classify fine-tuning-injected secrets, marking two representationally distinct regimes. We then introduce probe-geometry alignment (PGA), a surgical erasure that aligns activations along the probe's live readout direction at each depth. PGA drives the cross-sequence probe below random chance at all four scales tested (toy depth-4: 0.17; Pythia-70M: 0.07; Mistral-7B: 0.45; GPT-2 medium: 0.06 via MD-PGA k=2) and remains robust to six adversarial probe variants. Against a re-fitting attacker who trains a fresh probe on PGA-treated activations, we extend PGA adversarially, defeating the re-fit probe at every memorisation-relevant depth while preserving five zero-shot capability benchmarks within 2.8 percentage points per task (mean Δacc = +0.2pp). The cross-sequence signature is a real, causally separable, regime-specific property of pretrained representations -- removable below chance with a single rank-one intervention per depth at no measurable capability cost.
Sources
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- Scalable Extraction of Training Data from (Production) Language Models
- Who's Harry Potter? Approximate Unlearning in LLMs
- Large Language Model Unlearning
- Eight Methods to Evaluate Robust Unlearning in LLMs
- The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
- SoK: Machine Unlearning for Large Language Models
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Steering Language Models With Activation Engineering
- GPT-NeoX-20B: An Open-Source Autoregressive Language Model
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Mistral 7B
- Distilling the Knowledge in a Neural Network
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks