SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
cs.LG, cs.AI, cs.CL, cs.CR
Submitted: 2025-04-11
Updated: 2026-09-07
Comments: COLM 2025
Code: https://github.com/aashiqmuhamed/DynamicSAEGuardrails
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Machine unlearning is a promising approach to improve LLM safety by removing unwanted knowledge from the model.
Terminology
Abstract
Machine unlearning is a promising approach to improve LLM safety by removing unwanted knowledge from the model. However, prevailing gradient-based unlearning methods suffer from issues such as high computational costs, hyperparameter instability, poor sequential unlearning capability, vulnerability to relearning attacks, low data efficiency, and lack of interpretability. While Sparse Autoencoders are well-suited to improve these aspects by enabling targeted activation-based unlearning, prior approaches underperform gradient-based methods. This work demonstrates that, contrary to these earlier findings, SAEs can significantly improve unlearning when employed dynamically. We introduce Dynamic DAE Guardrails (DSG), a novel method for precision unlearning that leverages principled feature selection and a dynamic classifier. Our experiments show DSG substantially outperforms leading unlearning methods, achieving superior forget-utility trade-offs. DSG addresses key drawbacks of gradient-based approaches for unlearning -- offering enhanced computational efficiency and stability, robust performance in sequential unlearning, stronger resistance to relearning attacks, better data efficiency including zero-shot settings, and more interpretable unlearning.
Sources
- GPT-4 Technical Report
- Obfuscated Activations Bypass LLM Latent-Space Defenses
- Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
- Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Do Unlearning Methods Remove Information from Language Model Weights?
- Toy Models of Superposition
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
- Applying sparse autoencoders to unlearn knowledge in language models
- On Large Language Model Continual Unlearning
- Measuring Massive Multitask Language Understanding
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- An Adversarial Perspective on Machine Unlearning for AI Safety
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- Pointer Sentinel Mixture Models
- Steering Language Model Refusal with Sparse Autoencoders
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks