SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
cs.CL, cs.LG
Submitted: 2026-06-07
Updated: 2026-09-11
Comments: EMNLP 2026
Code: https://github.com/he-jingyi/SAEExplainer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a
Terminology
Abstract
Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Current explanation methods, however, typically operate within an open-loop paradigm, failing to leverage mechanistic feedback for further refinement. In this paper, we propose SAEExplainer, a training framework that utilizes activation scores as an objective reward signal to train the model for self-correction and iterative bootstrapping. By iteratively verifying and correcting foundational explanations through a two-round optimization process, SAEExplainer achieves continuous improvement in its explanatory capabilities. This mechanism significantly reduces explanation hallucinations and reinforces causal triggering patterns. Extensive experiments demonstrate our approach improves upon established baselines across most metrics, especially in causal triggering and discriminative activation. The code is available at https://github.com/he-jingyi/SAEExplainer.
Sources
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Towards Unifying Interpretability and Control: Evaluation via Intervention
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Gemma 2: Improving Open Language Models at a Practical Size
- OpenAI GPT-5 System Card
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- LatentQA: Teaching LLMs to Decode Activations Into Natural Language
- Automatically Interpreting Millions of Features in Large Language Models
- Improving Dictionary Learning with Gated Sparse Autoencoders
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering