Test-Time Unlearning via Sparse Autoencoder
cs.LG, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch.
Terminology
Abstract
Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.
Sources
- CRISP: Persistent Concept Unlearning via Sparse Autoencoders
- Steering When Necessary: Flexible Steering Large Language Models with Backtracking
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
- Who's Harry Potter? Approximate Unlearning in LLMs
- Applying sparse autoencoders to unlearn knowledge in language models
- Sparse Autoencoder Features for Classifications and Transferability
- Scaling and evaluating sparse autoencoders
- Gemma 3 Technical Report
- Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
- Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Rethinking Machine Unlearning for Large Language Models
- An Adversarial Perspective on Machine Unlearning for AI Safety
- Eight Methods to Evaluate Robust Unlearning in LLMs
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks