From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
summary
The gist
Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space, and this work introduces Mixture of Factor Analyzers
In short
This work introduces Mixture of Factor Analyzers (MFA) to decompose language model activations into local Gaussian regions. MFA models activation space as a collection of these regions, each defined by a centroid and a low-dimensional subspace capturing local variation. This provides an interpretable view of how concepts are realized in the model's internal representations.
Key concepts
- Mixture of Factor Analyzers (MFA)
- MFA models activation space as several Gaussian regions. Each region is defined by a mean centroid and a local factor analysis model that captures the structured variation within that specific area of the representation space.
- Centroid ($Ƶ_k$)
- The centroid represents the center or 'mean' of a specific Gaussian region in activation space. It defines the primary location or typical activation pattern for all data points belonging to that particular component.
- Local Variation ($z_{\hat{k}}$)
- This term describes the local variation captured by a component, modeled using a factor analysis subspace. It represents how activations shift away from the region's centroid, capturing fine-grained differences within that specific concept cluster.
Terminology used across episodes
This episode discusses
- From Directions to Regions: Decomposing Activations in Language Models via Local Geometry · Paper Radio
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Discovering Variable Binding Circuitry with Desiderata
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- The Llama 3 Herd of Models · Paper Radio
- When Models Manipulate Manifolds: The Geometry of a Counting Task
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
- The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
- GPT-4o System Card
- Automatically Interpreting Millions of Features in Large Language Models
- Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
- Open Problems in Mechanistic Interpretability
- OpenAI GPT-5 System Card
- Embedding Projector: Interactive Visualization and Interpretation of Embeddings
- Gemma 2: Improving Open Language Models at a Practical Size
- Steering Language Models With Activation Engineering
- Does BERT Make Any Sense? Interpretable Word Sense Disambiguation with Contextualized Embeddings
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
The paper
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry · Read on arXiv
Shafran Or Shafran Shaked Ronen Omri Fahn Shauli Ravfogel Atticus Geiger Mor Geva
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "From Directions to Regions".
Jane: Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and authors of this work, "From Directions to Regions: Decomposing Activations in Language Models via Local Geometry," because it really sets the stage for what they're proposing.
Jane: The paper is authored by Shafran Ronen, Omri Fahn, Shauli Ravfogel, Atticus Geiger, and Mor Geva; their work focuses on using Mixture of Factor Analyzers or MFA to decompose language model activations into these geometric regions.
Lu: What's interesting about the title is that it directly contrasts the old way of thinking with the new one; they are moving away from analyzing activations as just directions and toward understanding them as distinct regions defined by local geometry.
Meng: So, if I understand correctly, they are suggesting that instead of looking for one straight path to a concept, we should look at where that activation sits in a collection of overlapping shapes?
Lalam: It sounds like they are giving us a better map of the model's internal state; having these regions makes the representation space much more structured and manageable for us to work with.
The paper's summary: Tom: So, moving on to what they actually propose, the summary of this paper explains that MFA models activation space as a mixture of Gaussian regions, each defined by a mean centroid and its own local covariance structure.
Jane: Exactly; in simple terms, the core idea is that they take all those complex activations and partition them into groups—these are the Gaussian regions—and then for each group, they learn how it varies locally around its center point.
Lu: What's key here is that this approach allows them to capture complex, nonlinear structures in activation space, something traditional methods based on linear separability just miss.
Meng: So instead of treating every direction as equally important globally, the system identifies these specific clusters of activations that share similar local patterns and then models the variation inside those clusters separately.
Lalam: That sounds like a very sophisticated way to organize the information; it’s not just one big blob of data, but many localized areas with their own internal rules for how things shift.
The paper's improvements: Tom: Now let's discuss what they claim are the improvements this methodology offers over previous approaches, which is where it gets pretty compelling.
Jane: The main improvement is that MFA provides a way to decompose an activation into two distinct parts: a region in activation space and a within-region offset, which lets us analyze both the global position and the local variation separately.
Lu: They show that when you train these MFAs on large models like Llama-three point one-8B and Gemma-two-2B, they successfully capture complex patterns, including two distinct types of regions: narrow Gaussians for constrained lexical patterns and broad Gaussians for wide thematic topics.
Meng: That distinction between narrow and broad regions is interesting because it suggests that the model’s internal logic isn't uniform; some parts handle specific syntax while others manage very wide themes.
Lalam: And they mention that this decomposition leads to a significant increase in interpretability, achieving an average interpretability fraction of zero point nine six, which is much higher than what we usually see with other methods like Sparse Autoencoders.
Conclusion: Tom: So, wrapping up the discussion on "From Directions to Regions: Decomposing Activations in Language Models via Local Geometry," the paper proposes a way to move beyond simple global directions and into understanding activation space as a structured collection of regions with local geometry.
Jane: Essentially, they’ve shown that we can model these activations as a mixture of Gaussian distributions, giving us both the center point and the internal variation within each region for any given activation.
Lu: The implications are significant because this framework provides a scalable way to handle the complexity of modern language model representations without having to rely on assumptions about simple linear separability everywhere.
Meng: From a practical standpoint, if we can use these regions, we can get much better at steering the model's behavior because we can target either the broad theme of a region or the fine-grained local variation within it.
Lalam: I think this work is important because it gives us a way to quantify which parts of an activation are meaningful and interpretable, which is crucial for building AI systems that are more transparent and reliable in their operations.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck