From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

arXiv:2602.02464 · cs.CL · Submitted 2026-02-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "From Directions to Regions".

Jane: Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and authors of this work, "From Directions to Regions: Decomposing Activations in Language Models via Local Geometry," because it really sets the stage for what they're proposing.

Jane: The paper is authored by Shafran Ronen, Omri Fahn, Shauli Ravfogel, Atticus Geiger, and Mor Geva; their work focuses on using Mixture of Factor Analyzers or MFA to decompose language model activations into these geometric regions.

Lu: What's interesting about the title is that it directly contrasts the old way of thinking with the new one; they are moving away from analyzing activations as just directions and toward understanding them as distinct regions defined by local geometry.

Meng: So, if I understand correctly, they are suggesting that instead of looking for one straight path to a concept, we should look at where that activation sits in a collection of overlapping shapes?

Lalam: It sounds like they are giving us a better map of the model's internal state; having these regions makes the representation space much more structured and manageable for us to work with.

The paper's summary: Tom: So, moving on to what they actually propose, the summary of this paper explains that MFA models activation space as a mixture of Gaussian regions, each defined by a mean centroid and its own local covariance structure.

Jane: Exactly; in simple terms, the core idea is that they take all those complex activations and partition them into groups—these are the Gaussian regions—and then for each group, they learn how it varies locally around its center point.

Lu: What's key here is that this approach allows them to capture complex, nonlinear structures in activation space, something traditional methods based on linear separability just miss.

Meng: So instead of treating every direction as equally important globally, the system identifies these specific clusters of activations that share similar local patterns and then models the variation inside those clusters separately.

Lalam: That sounds like a very sophisticated way to organize the information; it’s not just one big blob of data, but many localized areas with their own internal rules for how things shift.

The paper's improvements: Tom: Now let's discuss what they claim are the improvements this methodology offers over previous approaches, which is where it gets pretty compelling.

Jane: The main improvement is that MFA provides a way to decompose an activation into two distinct parts: a region in activation space and a within-region offset, which lets us analyze both the global position and the local variation separately.

Lu: They show that when you train these MFAs on large models like Llama-three point one-8B and Gemma-two-2B, they successfully capture complex patterns, including two distinct types of regions: narrow Gaussians for constrained lexical patterns and broad Gaussians for wide thematic topics.

Meng: That distinction between narrow and broad regions is interesting because it suggests that the model’s internal logic isn't uniform; some parts handle specific syntax while others manage very wide themes.

Lalam: And they mention that this decomposition leads to a significant increase in interpretability, achieving an average interpretability fraction of zero point nine six, which is much higher than what we usually see with other methods like Sparse Autoencoders.

Conclusion: Tom: So, wrapping up the discussion on "From Directions to Regions: Decomposing Activations in Language Models via Local Geometry," the paper proposes a way to move beyond simple global directions and into understanding activation space as a structured collection of regions with local geometry.

Jane: Essentially, they’ve shown that we can model these activations as a mixture of Gaussian distributions, giving us both the center point and the internal variation within each region for any given activation.

Lu: The implications are significant because this framework provides a scalable way to handle the complexity of modern language model representations without having to rely on assumptions about simple linear separability everywhere.

Meng: From a practical standpoint, if we can use these regions, we can get much better at steering the model's behavior because we can target either the broad theme of a region or the fine-grained local variation within it.

Lalam: I think this work is important because it gives us a way to quantify which parts of an activation are meaningful and interpretable, which is crucial for building AI systems that are more transparent and reliable in their operations.

Shafran Or Shafran Shaked Ronen Omri Fahn Shauli Ravfogel Atticus Geiger Mor Geva

cs.CL

Submitted: 2026-02-02

Updated: 2026-09-28

Comments: Accepted at ICML 2026 main conference

Code: https://github.com/ordavid-s/decomposing-activations-local-geometry

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space, and this work introduces Mixture of Factor Analyzers

Key concepts

Mixture of Factor Analyzers (MFA)
MFA models activation space as several Gaussian regions. Each region is defined by a mean centroid and a local factor analysis model that captures the structured variation within that specific area of the representation space.
Centroid ($Ƶ_k$)
The centroid represents the center or 'mean' of a specific Gaussian region in activation space. It defines the primary location or typical activation pattern for all data points belonging to that particular component.
Local Variation ($z_{\hat{k}}$)
This term describes the local variation captured by a component, modeled using a factor analysis subspace. It represents how activations shift away from the region's centroid, capturing fine-grained differences within that specific concept cluster.

Terminology

Summary

Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space, and this work introduces Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models activation space as a collection of Gaussian regions with their local covariance structure. MFA decomposes activations into two compositional geometric objects: the region’s centroid in activation space, and the local variation from the centroid.

The gist

MFA maps the activation space into Gaussian regions, where each region is modeled with a low-dimensional subspace that captures structured within-region variation, thereby providing an interpretable decomposition of language model activations.

How it works

  1. MFA models the space as a collection of local low-dimensional Factor Analyzers (FA), where each component models a different region of the representation space.

  2. Each component k is defined by a mean centroid µk, which sets the center of its region, and a factor analysis model with a component-specific matrix Wk, which determines the orientation of its local low-dimensional subspace.

  3. The generative model for an activation x conditioned on component k is given by x = µk + Wkzk + ϵ, where z is a latent factor vector and ϵ is noise with a diagonal covariance matrix Ψ.

  4. The overall density of the activation space is represented as a mixture of these component densities: p(x) = Σ πk N (x µk, Ck), where πk are the mixture weights.

Decomposition and Interpretation

The decomposition yields two compositional geometric objects for any given activation x: a region in activation space defined by the assigned centroid (Rk(x) µk), and a within-region offset defined by the local variation captured by the component's subspace (Rk(x) zˆk). The reconstruction of an activation is then written as a linear product: x ≈ A b(x), where A is formed by concatenating the component means and loadings, and b(x) concatenates the responsibilities and latent coordinates. This structure moves the unit of analysis from isolated global directions to local regions with their own low-rank geometry.

Structural Analysis of Discovered Regions

Training large-scale MFAs on Llama-3.1-8B and Gemma-2-2B reveals rich structures in the activation space, characterized by two classes of Gaussians: narrow Gaussians that concentrate on a constrained lexical pattern (showing more syntactic variance) and broad Gaussians that encompass wide thematic topics (often exhibiting semantic local variation). Analyzing these regions shows that neighboring components tend to encode related semantics, suggesting concepts may be realized by constellations of nearby Gaussians. Furthermore, the analysis of loadings reveals that within-region variation often reflects both semantic and syntactic differences, with narrow Gaussians skewing more syntactic and broad Gaussians skewing more semantic.

Performance on Localization and Steering

MFA is evaluated on localization (MCQA and RAVEL) and steering benchmarks. On localization, MFA outperforms large-scale SAEs by large margins, beating supervised baselines like Desiderata-Based Masking (DBM) in several tasks. For causal steering, utilizing the MFA centroids steers better than SAE features in the majority of settings, typically exhibiting a twofold gain on coherence and conceptual alignment. The results indicate that absolute positions learned by the centroids are an effective unit for steering, while local offsets capture fine-grained shifts within the broader theme. MFA achieves an average interpretability fraction (IF) of 0.96 ± 0.2, significantly higher than SAEs' average IF of 0.29 ± 0.2, indicating that most of the high-contribution features in its decomposition are interpretable.

Reconstruction and Limitations

When comparing reconstruction error on a held-out validation set, MFA substantially outperforms the K-means baseline, which consistently has 1.3 − 1.5 times higher MSE, demonstrating that modeling within-region variation with a learned low-rank structure captures a considerable portion of the structured representation. However, a key limitation is that MFA explicitly models the activation distribution it is trained on. This means that when an activation lies in a region that is rare or out of distribution of the training set, MFA may assign it to the nearest available component even if none provides a good local fit, leading to high reconstruction error. Despite this, MFA isolates meaningful and useful features that generalize to tasks like steering and localization.

Conclusion

MFA proposes a local-geometry view of activation space, partitioning it into low-dimensional regions and modeling the intrinsic modes of variation within each region, offering a scalable approach to model control that generalizes across layers and models.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the concepts from this scientific paper, and what those improved systems will be able to do:


The core improvement lies in shifting from analyzing activations as isolated global directions (the limitation of Sparse Autoencoders or simple Global Direction methods) to analyzing them as a structured collection of local regions defined by low-dimensional Gaussian distributions.

Here are the specific improvements and resulting capabilities:

  1. Implementation of Mixture of Factor Analyzers (MFA) for Activation Decomposition:

  2. Improved Conceptual Disentanglement via Region Assignment and Local Subspace Analysis:

  3. Enhanced Causal Localization Capabilities through Centroid Interventions:

  4. Superior Controllability and Steering via Local Variation Manipulation:

Specific Improvements and Capabilities of the Improved AI System:

  1. Implementation of MFA for Activation Decomposition (Regions + Local Subspaces):

  2. The system will decompose any activation into two compositional geometric objects: a region (its centroid in activation space) and a within-region offset (the local variation parameterized by a low-dimensional subspace).

  3. Improved Conceptual Disentanglement via Region Assignment and Local Subspace Analysis:

  4. The AI will no longer rely on isolated global directions to represent concepts. Instead, it will discover clusters of nearby Gaussian regions, where neighboring Gaussians collectively form a unified semantic neighborhood (e.g., grouping diverse movie genres under one broad Gaussian).

  5. Enhanced Causal Localization Capabilities through Centroid Interventions:

  6. The system can perform precise causal localization by manipulating the region's centroid (absolute position in activation space). Intervening on the centroid reliably steers the model toward a broad, high-level semantic theme associated with that entire region (e.g., steering toward Sports by moving to the Sports Gaussian centroid).

  7. Superior Controllability and Steering via Local Variation Manipulation:

  8. The system gains fine-grained control over sub-concepts within a broad theme by intervening on the local subspace directions (the loadings, represented as within-region offsets). This allows for targeted shifts to specific subthemes or syntactic variations even when operating within a broad conceptual region (e.g., steering toward Basketball from the general Sports Gaussian centroid).

  9. Improved Interpretability and Feature Relevance:

  10. The system provides a quantifiable measure of feature interpretability (IF score, averaging 0.96 in the paper), ensuring that the features it isolates are coherent and meaningful to human annotators, rather than being artifacts of the training objective.

In summary, this research enables an AI system to move from simply identifying what a concept is represented by (a single direction) to understanding where a concept lives in the activation space (the region/centroid) and how it manifests internally (the local geometry/subspace), leading to systems that are more robust, controllable, and transparent for complex language understanding tasks.

Abstract

Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for individual global directions, implicitly assuming linear separability, which overlooks concepts with nonlinear or multi-dimensional structure. In this work, we leverage Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models the activation space as a collection of Gaussian regions with their local covariance structure. MFA decomposes activations into two compositional geometric objects: the region's centroid in activation space, and the local variation from the centroid. We train large-scale MFAs for Llama-3.1-8B and Gemma-2-2B, and show they capture complex, nonlinear structures in activation space. Moreover, evaluations on localization and steering benchmarks show that MFA outperforms unsupervised baselines, is competitive with supervised localization methods, and often achieves stronger steering performance than sparse autoencoders. Together, our findings position local geometry, expressed through subspaces, as a promising unit of analysis for scalable concept discovery and model control, accounting for complex structures that isolated directions fail to capture.

Sources

Related papers