Exemplar Partitioning for Mechanistic Interpretability
cs.LG
Submitted: 2026-05-14
Updated: 2026-09-17
Comments: Code: https://github.com/jessicarumbelow/exemplar-partitioning. Pretrained dictionaries: https://huggingface.co/datasets/J-RUM/exemplar-partitioning
Code: https://github.com/jessicarumbelow/exemplar-partitioning
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Exemplar Partitioning (EP), an unsupervised method for building interpretable feature dictionaries from large language model activations.
Terminology
Abstract
We introduce Exemplar Partitioning (EP), an unsupervised method for building interpretable feature dictionaries from large language model activations. An EP dictionary is a Voronoi partition of activation space, built by leader-clustering streamed activations within a distance threshold. Each region is defined by an observed exemplar and an average of its member activations, which define region membership and provide directions for intervention. Dictionary size is determined by the activation stream at the chosen threshold rather than pre-selected. Exemplars link regions to observed inputs, allowing dictionaries built from the same input stream to be compared across layers, training checkpoints, and architectures. We demonstrate how EP can be used to interpret and intervene on model behaviour, track changes in activation space through training, and detect hidden concepts. Comparing EP dictionaries on base and instruction-tuned Gemma-2-2B and Llama-3.1-8B reveals that instruction tuning reorganises harmful prompt activations similarly across the two models, but at different granularities. Interventions on these regions make both models answer harmful requests they previously refused. In 19 of 21 Taboo models trained to hide a secret word, EP finds new regions that do not exist in the base model, whose decoded tokens relate to the known secret on inspection. Although EP assigns each token to a single region, linear probes built from EP regions achieve up to 90.5% of full-activation probe accuracy. On AxBench concept detection at Gemma-2-2B-it layer 20, EP achieves the highest mean AUROC of all unsupervised methods (0.937), outperforming SAE-A (0.911) and approaching supervised probes (0.946). Building EP dictionaries is fast and cheap: EP uses about 10 cubed times fewer construction tokens than comparable SAEs.
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Scaling and evaluating sparse autoencoders
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Steering Llama 2 via Contrastive Activation Addition
- Sparse Autoencoders Trained on the Same Data Learn Different Features
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Attribution Patching Outperforms Automated Circuit Discovery
- Steering Language Models With Activation Engineering
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks