Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
Zirui Song, Huaxing Liu, Xiang Wang, Shuai Li, Xinye Li, Lang Gao, Jinghui Zhang, Zheng Lu, Fengxian Ji, Xiaojun Chang, Xiuying Chen
Mohamed bin Zayed University of Artificial Intelligence · AMAP, Alibaba Group · Peking University
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: In processing
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper introduces J-Access, an inference-time auditing method that uses the Jacobian lens to map intermediate representations into vocabulary space and measure how often target concepts remain
Terminology
Summary
The paper introduces J-Access, an inference-time auditing method that uses the Jacobian lens to map intermediate representations into vocabulary space and measure how often target concepts remain accessible along a model's output pathway. The authors propose this as a diagnostic tool for assessing residual knowledge accessibility in LLMs after machine unlearning.
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued training or serve as reliable optimization targets.
The paper addresses two practical questions: (1) Can residual signals predict future knowledge recovery under continued training? (2) Do these signals remain reliable when directly optimized?
The J-Access audit uses the Jacobian lens to transport intermediate representations through a linear approximation of the model's downstream computation before decoding them into vocabulary space. For each probe query that implicates a forgotten entity without naming it, the method maps mid-depth residual representations from a band of mid-to-late layers into vocabulary space and checks whether any token associated with the target concept appears among the top-ranked decoded tokens. The resulting access rate is normalized between the original model and a retain-only gold model.
The score is defined as:
- JOcc(θ) = (Ak(θ; I(θ)) − Ak(θg; I(θ))) / (Ak(θ0; I(θ)) − Ak(θg; I(θ)))
where θ0 is the original model, θg is the gold retain-only model, and I(θ) is the set of probes on which the model does not behaviorally produce the target concept.
The authors evaluate on the TOFU benchmark and OpenUnlearning, spanning eight unlearning methods (GradDiff, NPO, SimNPO, AltPO, IdkDPO, IdkNLL, UNDIAL, RMU) and 398 public unlearned models.
Across all 398 unlearned models, 85% exhibit higher target access than the retain-only gold model, and the median normalized accessibility is 0.69. The typical unlearned model closes less than one third of the measured internal gap between the original and gold models.
Models with similar behavioral scores span a wide range of accessibility values, both across and within method families. J-Access agrees with the independent activation-patching criterion (Unlearning DepthScore) in the expected direction: unlearned models with greater residual access exhibit shallower causal deletion (Spearman ρ = −0.72).
Higher pre-attack J-Access is associated with greater recovery under relearning attacks:
-
Full pool correlation with excess revival: ρ = +0.35
-
Partial correlation controlling for Forget Quality and Model Utility: ρ = +0.33
-
Within-method pooled correlation: ρ = +0.45 (identity-level), increasing to ρ = +0.71 for the knowledge-level variant
-
Steps-to-recover correlation: ρ = −0.70 (higher accessibility requires fewer fine-tuning steps to recover)
The relationship holds within each of the eight unlearning methods, indicating it captures checkpoint-level variation rather than merely identifying weaker algorithms.
However, item-level prediction remains near chance: J-Access does not reliably distinguish which specific facts will recover (AUROC ≈ 0.5), and adding it to behavioral and membership-inference predictors yields negligible improvement. The authors interpret this as consistent with prior localization studies suggesting fine-tuning-based unlearning disables a shared retrieval pathway rather than erasing individual stored facts.
Directly minimizing J-Access during unlearning (via WD-Train with suppression weights λ ∈ 0, 5, 10) produces the opposite of the expected pattern:
-
Increasing the suppression weight from λ = 0 to λ = 10 lowers J-Access from 0.67 to 0.55
-
But increases post-attack revival from 0.283 to 0.387
-
UDS (causal deletion depth) remains nearly unchanged
In contrast, models with deeper causal deletion under UDS consistently exhibit greater resistance to relearning: GradDiff and RMU configurations with high UDS achieve post-attack revival rates of only 0.025 and 0.000, respectively, whereas their shallow-deletion counterparts reach 0.800 and 0.695.
The authors conclude: "Optimizing J-Access suppresses evidence exposed to the audit without removing the knowledge that supports recovery. This asymmetry shows that a useful independent diagnostic is not necessarily a valid unlearning objective."
The paper argues that internal audits should serve as an independent diagnostic dimension in unlearning evaluation, and should not be converted into optimization targets without validation.
Reliable unlearning evaluation requires internal audits to complement behavioral metrics and be validated against independent causal and recovery evidence, rather than treated as deletion certificates or optimization targets.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Add a
residual knowledge accessibility
diagnostic to unlearning pipelines. AI systems that undergo machine unlearning (e.g., removing copyrighted or private data) can be equipped with a J-Access-style audit layer that maps intermediate representations into vocabulary space. This allows the system to report a normalized accessibility score (0–1) alongside behavioral forgetting metrics, giving operators a clear signal of whether the model has truly erased knowledge or merely suppressed output. -
Implement a two-stage unlearning validation protocol. Improved AI systems will first run behavioral tests (does the model refuse to output the target?) and then run a J-Access audit (is the target still internally accessible?). If J-Access > 0.3, the system flags the unlearning as incomplete and recommends further fine-tuning or a different method—preventing false confidence in
forgotten
models. -
Use J-Access as a pre-attack risk predictor for relearning attacks. AI systems can be trained to estimate vulnerability to future knowledge recovery by computing J-Access before any adversarial fine-tuning. A high J-Access score (e.g., > 0.7) would trigger automated safeguards, such as adding noise to residual streams or applying stronger regularization, reducing the risk of knowledge revival during continued training.
-
Avoid using J-Access as an optimization target. The improved AI system will explicitly reject any training loop that minimizes J-Access directly, since the paper shows this suppresses audit evidence without removing underlying knowledge (revival increases from 0.283 to 0.387). Instead, the system will use J-Access only for monitoring and validation, never as a loss function component.
-
Enhance unlearning method selection via causal-depth correlation. The system can use J-Access in conjunction with causal deletion depth (UDS) to rank unlearning methods. For example, it can automatically prefer configurations where J-Access is low and UDS is high (e.g., GradDiff with deep deletion, revival rate 0.025) over those with low J-Access but shallow deletion (revival rate 0.800), improving long-term robustness.
-
Provide item-level uncertainty warnings. Since J-Access fails to predict which specific facts will recover (AUROC ≈ 0.5), the improved AI system will not claim per-fact deletion guarantees. Instead, it will output aggregate risk scores with confidence intervals, and warn users that individual fact recovery is unpredictable—encouraging conservative deployment for sensitive data.
-
Enable continuous internal audit logging. The system can periodically recompute J-Access during fine-tuning or after updates, tracking how accessibility evolves. If J-Access rises above a threshold post-deployment, the system can alert administrators to potential knowledge leakage, enabling proactive re-unlearning.
What the improved AI system can do: It can perform unlearning with verifiable internal erasure, predict its own vulnerability to knowledge revival, reject unsafe optimization targets, select robust unlearning methods, and provide honest, uncertainty-aware reporting—all while avoiding the false security of behavioral-only evaluations.
Abstract
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued training or serve as reliable optimization targets. Resolving this gap is essential to determine whether internal auditing can move beyond post-hoc evaluation toward proactive risk monitoring and safer unlearning. We propose J-Access, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway. We hypothesize that residual accessibility reflects recovery susceptibility: knowledge that remains closer to the output pathway requires less fine-tuning to restore, leading to faster recovery. We audit 398 public unlearned models spanning eight unlearning methods. We find that: (1) most unlearned models retain access above the retain-only gold level; (2) pre-attack accessibility predicts recovery speed and extent at the model level, but cannot identify which specific facts will be recovered; and (3) directly minimizing J-Access does not promote genuine deletion. Instead, the model learns to hide knowledge from the audit, producing lower audit scores but greater post-attack recovery. These findings position J-Access as a model-level diagnostic for assessing residual susceptibility in unlearned models. We argue internal audits should serve as an independent diagnostic dimension in unlearning evaluation, and should not be converted into optimization targets without validation.
Sources
- Obfuscated Activations Bypass LLM Latent-Space Defenses
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Machine Unlearning
- Discovering Latent Knowledge in Language Models Without Supervision
- Do Unlearning Methods Remove Information from Language Model Weights?
- Who's Harry Potter? Approximate Unlearning in LLMs
- Scaling Laws for Reward Model Overoptimization
- Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
- Certified Data Removal from Machine Learning Models
- Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
- Verbalizable Representations Form a Global Workspace in Language Models
- Measuring the Depth of LLM Unlearning via Activation Patching
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- An Adversarial Perspective on Machine Unlearning for AI Safety
- Eight Methods to Evaluate Robust Unlearning in LLMs
- Towards eliciting latent knowledge from LLMs with mechanistic interpretability
- Extracting Unlearned Information from LLMs with Activation Steering
- UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
- Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering