SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Weihan Meng, Hongzhu Guo, Yi Jing, Dewen Liu, Zijun Yao, Xiaozhi Wang, Lei Hou, Juanzi Li
Tsinghua University · Peking University · Fudan University
cs.CL
Submitted: 2026-08-18
Updated: 2026-08-19
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization Summary This paper introduces SAEVerbalizer, a framework that fine-tunes large language models
Terminology
Summary
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Summary
This paper introduces SAEVerbalizer, a framework that fine-tunes large language models (LLMs) to generate natural-language explanations of sparse autoencoder (SAE) features directly from their decoder directions, moving beyond traditional methods that rely on external observation of model behavior.
Problem and Motivation
Sparse autoencoders (SAEs) are widely used to interpret LLM internal representations by mapping dense representations into a higher-dimensional sparse feature space, where each dimension represents an interpretable feature. However, while SAEs excel at extracting features, they fall short of explaining them in natural language.
Existing methods explain SAE features through externally observed LLM behavior, typically by running an LLM over a large corpus, identifying each feature's top-activating examples, and prompting an LLM to summarize their shared pattern. These methods face two limitations: (1) Superficial Explanations—without directly examining how a feature is represented inside the LLM, explanations describe what activation examples have in common rather than what the feature itself represents; (2) Computational Inefficiency—each new SAE requires repeating corpus-scale inference, activation computation, example retrieval, and feature-wise LLM summarization.
Method
SAEVerbalizer shifts SAE feature explanation from observing external behavior to processing internal representations within the LLM.
The framework injects an SAE decoder direction—the only feature-specific input—into the LLM's representations and prompts the LLM to generate a natural-language explanation of the corresponding feature.
Verbalizer Design: The verbalizer receives a fixed, feature-agnostic prompt specifying the verbalization task. During prompt prefilling, the target decoder direction is injected at a chosen layer (Linj) into each token representation in a designated injection span. The injection uses norm-matched additive injection: h′b,s = hb,s + αn̄b v̂fb, where v̂fb is the normalized decoder direction, n̄b is the mean pre-injection representation norm over the injection span, and α controls relative strength. No activation examples or surrounding contexts are provided. The remaining Transformer layers process the post-injection prompt representations, conditioning subsequent autoregressive generation to yield an explanation.
Adapter Design: For SAE features from another LLM, a lightweight adapter maps source-layer representation spaces to the verbalizer's injection-layer representation space. The adapter is parameterized as a single affine layer: A(x) = Wx + b. At inference, the mapped direction is computed as the change the source decoder direction induces in the adapter output: A(h(s) + vf(s)) − A(h(s)) = Wvf(s). The affine bias cancels in this difference.
Training: The verbalizer is fine-tuned on high-quality feature–explanation pairs obtained from Neuronpedia, filtered through a two-stage LLM-based filtering procedure scoring coherence, specificity, and consistency. The training objective is standard causal language modeling with token-level cross-entropy computed only over explanation tokens. Partial fine-tuning freezes layers up to and including Linj, trains layers after Linj, and freezes the embedding layer while fine-tuning the final normalization layer and language modeling head. The adapter is trained on aligned hidden representations from unlabeled text using mean squared error loss, with both LLMs frozen.
Experiments and Results
The paper evaluates verbalizers fine-tuned on gemma-3-1b-it, gemma-3-4b-it, and gemma-3-27b-it backbones using Gemma Scope 2 SAEs (width-262k, medium sparsity). Evaluation uses Reference Agreement (RA), measuring the proportion of test features for which generated explanations agree with Neuronpedia references as judged by an LLM judge.
Main Results: Verbalizers generalize to unseen features across all backbone–layer configurations. The best configuration (27B-L16 with 48k training pairs) achieves RA scores of 52.3%, 80.5%, and 56.1% on Global Train-Standard (GTS), Low-Index Gold (LIG), and Global Gold (GG) test sets, respectively. RA generally increases with backbone scale and tends to be higher at earlier layers within each backbone.
Transferability: The verbalizer transfers across SAE dictionaries directly (without adapter or supervision) when SAEs share the same representation space—the default verbalizer achieves substantial agreement on an unseen width-65k SAE (64.4% GTS, 56.5% LIG, 65.9% GG). Cross-LLM transfer uses separately trained adapters: for 1B-L7 features, adapter-based transfer improves over the native verbalizer (18.2% vs. 17.6% GTS, 39.0% vs. 35.5% LIG, 17.4% vs. 14.7% GG), demonstrating that cross-LLM adaptation can leverage the stronger verbalization capability of the 27B verbalizer. For 4B-L9, transfer does not yield gains (32.8% vs. 37.6% GTS).
Ablation Study:
-
Supervision scaling: Fine-tuning on just 1.5k pairs yields a large gain over the zero-supervision backbone (36.4% vs. 1.6% GTS). Additional supervision improves GTS and GG, while LIG saturates early.
-
Prompt and injection robustness: Semantically similar prompts, nearby injection spans, and additive versus interpolative injection produce only minor differences. Performance is broadly stable across injection strengths (α from 0.10 to 1.00), but degrades substantially when the injected direction becomes too weak (α = 0.01 yields 18.9% GTS vs. 52.3% at default α = 0.2).
Case Studies:
-
Qualitative comparisons show the verbalizer identifies localized lexical or structural patterns or different semantic abstractions compared to Neuronpedia explanations (e.g., feature #14949: Neuronpedia says
lasting, enduring, eternal states
while verbalizer saysforever
; feature #115968: Neuronpedia saysGu followed by letters or syllables
while verbalizer saysgu- prefix
). -
Joint feature injection produces explanations combining meanings of both features (e.g., #15187
love and enthusiasm
+ #1586coffee and its contexts
→love of coffee
). -
Direction reversal produces semantically related but shifted meanings (e.g., #40105
lessons learned
reversed →lesson plan
).
Conclusion and Limitations
The paper concludes that "internal representation verbalization [is] a trainable and partially reusable complement to methods that infer SAE feature meanings from activation examples, providing a more direct route from learned representations to feature explanations." Limitations include: evaluation primarily on Gemma LLMs and Gemma Scope 2 SAEs; Reference Agreement measures consistency with filtered references rather than absolute correctness; each configuration is evaluated from a single run; and the supervision pipeline relies on computationally expensive filtered Neuronpedia explanations.
Improvements for AI systems
Improvements to AI Systems:
- Direct Feature-to-Language Translation for Interpretability
-
Improved system: An AI that can take any internal feature vector (from SAEs or similar sparse representations) and generate a natural-language explanation without needing to run inference over large corpora or retrieve activation examples.
-
Capability: Real-time, on-demand interpretation of individual neurons or feature directions during model operation, enabling faster debugging and auditing of AI behavior.
- Cross-Model Feature Transfer with Lightweight Adapters
-
Improved system: A universal interpretability layer that maps feature directions from one LLM’s representation space to another’s using a single affine transformation.
-
Capability: A small (e.g., 7B) model can leverage the explanatory power of a larger (e.g., 27B) model to interpret its own features, reducing the need for large-scale interpretability infrastructure per model.
- Zero-Shot Feature Composition and Manipulation
-
Improved system: An AI that can combine multiple feature directions (via additive injection) to generate explanations of conjunctive concepts (e.g.,
love of coffee
) and reverse feature directions to explore semantic opposites or shifts. -
Capability: Interactive exploration of feature spaces—users can probe how features interact, negate, or blend, enabling hypothesis generation about model reasoning.
- Scalable, Corpus-Free Interpretability Pipeline
-
Improved system: A self-contained interpretability module that replaces corpus-scanning and example-retrieval steps with direct representation processing.
-
Capability: Interpretability for new SAEs or model updates in minutes (not hours/days), making it feasible to audit models continuously during training or deployment.
- Layer-Aware Explanation Generation
-
Improved system: An AI that adapts its explanation style and abstraction level based on the source layer of the feature (e.g., earlier layers → more concrete/lexical, later layers → more abstract/semantic).
-
Capability: More nuanced interpretability reports that reflect the hierarchical nature of representations, helping researchers understand how concepts form across depth.
- Robustness to Injection Strength and Prompt Variation
-
Improved system: An AI that maintains explanation quality across a wide range of injection strengths (α = 0.1–1.0) and semantically equivalent prompts, making it reliable in production where exact hyperparameters may vary.
-
Capability: Stable interpretability outputs even when integrated into dynamic pipelines or used by non-expert operators.
- Supervision-Efficient Fine-Tuning for New Domains
-
Improved system: A verbalizer that achieves strong performance with as few as 1.5k training pairs (36.4% RA vs. 1.6% baseline), enabling rapid adaptation to new model families or specialized domains (e.g., medical, legal) with minimal labeled data.
-
Capability: Quick deployment of interpretability tools for niche or proprietary models without requiring massive annotation efforts.
- Joint Feature Injection for Multi-Concept Reasoning
-
Improved system: An AI that can accept multiple feature directions simultaneously and generate explanations that capture their intersection or interaction.
-
Capability: Advanced debugging of compositional reasoning—e.g., identifying when a model conflates or combines concepts, or exploring how features co-activate in complex tasks.
- Direction Reversal for Counterfactual Interpretability
-
Improved system: An AI that can explain what a feature negates or opposes by reversing the injected direction.
-
Capability: Provides contrastive explanations (e.g.,
lessons learned
vs.lesson plan
), helping users understand what a feature is not encoding, which is critical for bias detection and safety analysis.
- Unified Interpretability Across SAE Sizes and Sparsities
-
Improved system: A verbalizer that transfers directly across SAE dictionaries within the same representation space (e.g., from width-262k to width-65k) without retraining.
-
Capability: A single interpretability model can serve multiple SAE configurations, reducing maintenance overhead and enabling consistent explanations across different granularities of feature extraction.
Abstract
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an LLM's representations and fine-tunes the LLM's downstream layers to generate natural-language explanations of the injected features. Once trained, the resulting verbalizer explains SAE features directly from decoder directions, addressing both limitations. Our experiments show that the learned verbalization capability generalizes to unseen features, transfers across separately trained SAE dictionaries, and, with a lightweight adapter, extends to SAE features from different LLMs. Intervention experiments show that injecting multiple directions yields an explanation combining their meanings, while reversing individual directions produces corresponding meaning shifts.
Sources
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Gemma 3 Technical Report
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
- Training Language Models to Explain Their Own Computations
- Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering