When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/osu-srml/sae-robustness-under-pruning
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood.
Terminology
Abstract
Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation-aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer-wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings. Code is publicly available at https://github.com/osu-srml/sae-robustness-under-pruning/tree/main.
Sources
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Toy Models of Superposition
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- SlimLLM: Accurate Structured Pruning for Large Language Models
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Stable and Steerable Sparse Autoencoders with Weight Regularization
- Mistral 7B
- Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- LLM-Pruner: On the Structural Pruning of Large Language Models
- Compact Language Models via Pruning and Knowledge Distillation
- A Simple and Effective Pruning Approach for Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks