Measuring Semantic Abstractness of SAE Features via Nonlocality
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
University of Oxford · Stanford University
cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 18 pages, 9 figures
Code: https://github.com/lccqqqqq/sae-feature-nonlocality
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper addresses a central challenge in Mechanistic Interpretability (MI): distinguishing surface-level, token-driven features from genuinely high-level, abstract features in Sparse Autoencoders
Terminology
Summary
The paper addresses a central challenge in Mechanistic Interpretability (MI): distinguishing surface-level, token-driven features from genuinely high-level, abstract features in Sparse Autoencoders (SAEs). While SAEs have helped uncover mechanistic explanations for LLM behaviors such as reasoning and jailbreaking, existing feature selection methods—including keyword filtering, contrastive dataset curation, and auto-interpretation—do not adequately distinguish mechanistic complexity of features. The authors note that successful control of a target behaviour by intervening on a feature does not imply understanding of the underlying mechanism,
citing examples where steering on a feature might simply boost token probabilities (e.g., Wait
) rather than represent genuine uncertainty detection.
The paper introduces Feature Nonlocality (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation.
Formally:
For a feature a at layer l with activation z a(T, P) at position T of prompt P, the per-position influence is computed as the squared gradient norm:
J a(t, T, P) = ‖∂z a(T, P)/∂x t‖2
where x t are token representations at the layer-0 residual stream. This is normalized to a probability distribution over prefix positions, and the per-prompt nonlocality is the Shannon entropy:
H(a, P) = −Σ p(P) a*(t) log2 p(P) a*(t)
The dataset-level FNL averages this over firing events. The intuition is that Token-level features (e.g. bigram statistics) depend on local prefix cues and have low nonlocality, while abstract, topical features draw on contextual evidence spread broadly across the window and have high nonlocality.
The definition draws inspiration from holographic duality in theoretical physics.
FNL rankings are stable across topically diverse corpora (WikiText, GSM8K, Code-Python). For Gemma-2-2B, Spearman correlations between corpora range from 0.710 to 0.918 across layers 5, 12, and 20. The supplementary material extends this to Llama-3-8B and Qwen3-8B, reporting disattenuated correlations (ρ̃) of 0.531–0.981, showing the ranking retains the majority of its attainable agreement at every model and depth.
FNL rises over the first half of the network and saturates at middle layers (l ≈ 13–16 for Gemma-2-2B), broadly following the corresponding residual-stream statistic.
This pattern is robust across models and corpora.
FNL varies smoothly with decoder cosine similarity: features close together under decoder cosine similarity occupy similar ranges of nonlocality.
The supplementary material shows this correlation is with write-side geometry rather than semantic similarity—semantic neighbors (by auto-interpretation embedding cosine) spread widely in FNL, while decoder neighbors cluster narrowly.
Using the methodology from Ma et al. (2026), the authors test 100 reasoning features in DeepSeek-R1-Distill-Llama-8B. FNL correlates negatively with activation recovery under token injection (Spearman ρ = −0.39 to −0.46 across layers 10–12 and 19), and AUC values of 0.73–0.84 for distinguishing token-driven from context-dependent features. The mean FNL for context-dependent features (4.57–4.71 bits) consistently exceeds that for token-driven features (4.03–4.38 bits).
FNL correlates positively with paraphrase robustness score S a (ρ = 0.27, p = 0.011, n = 96 features), where S a measures retention of activation under meaning-preserving paraphrases versus token shuffles.
Representative features show a clear continuum: single-token detectors (e.g., f608 Robert indicator,
H(a) = 0.98 bits) sit at the low end, while epistemic hedging features (e.g., f13603, H(a) = 5.09 bits) sit at the high end. The paper concludes FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature.
The authors reproduce the CC-Delta pipeline (Assogba et al. 2026) for jailbreak-mitigation features in DeepSeek-R1-Distill-Llama-8B. Strikingly, 21 of the 25 selected features have 0 nonlocality
—they activate almost exclusively at position-0 of plain harmful prompts, functioning as positional/BOS indicators rather than content recognizers. Steering experiments confirm: the positional subset (21 features) achieves held-out safety of 0.911 versus 0.504 for the content subset (4 features), with the unsteered baseline at 0.511. This reveals the features should be understood as surface-level, beginning-of-sequence (BOS) indicators rather than encoding the computation for recognizing harmful intents.
For DeepSeek-R1-Distill-Llama-8B at layer 19, steering the top 20% FNL features improves MATH-500 accuracy by 4.6 points over unsteered baseline (0.911 vs. 0.865), outperforming low-FNL steering (+3.8 points), random features (+3.6 points), and a representative ReasonScore-selected feature (+3.9 points). However, the authors caution this is a proof of concept rather than evidence that FNL reliably predicts steering utility
—supplementary experiments show that outside DeepSeek-R1-Distill-Llama-8B, both high and low FNL arms fall below baselines.
-
FNL is a stable, label-free, gradient-based metric requiring
no contrastive datasets or interventions.
-
FNL correlates with existing proxy measures of semantic abstractness (token-injection recovery and paraphrase robustness).
-
FNL successfully distinguishes context-dependent reasoning features from token-driven ones in 73–84% of random pairs.
-
FNL provides diagnostic value for safety-relevant behaviors, revealing that effective jailbreak-mitigation features are often positional rather than content-aware.
-
FNL shows promise as a feature selection criterion for steering, though gains are model-specific.
The paper concludes that FNL helps distinguish what an intervention achieves and what the intervened feature represents
—the former validated by steering utility, the latter requiring additional evidence that FNL supplies.
Improvements for AI systems
Improvement 1: Safety-Feature Auditing for Jailbreak Mitigation
-
What I can do: Automatically screen any SAE feature selected for safety interventions (e.g., refusal, harmlessness, or jailbreak-mitigation features) by computing its FNL score. If a feature has near-zero nonlocality (e.g., activates only at position 0 or on BOS tokens), I can flag it as a positional artifact rather than a content recognizer.
-
Improved system capability: An AI safety pipeline that, before deploying a steering vector or activation patch, rejects features with FNL below a threshold (e.g., < 0.5 bits) to avoid false confidence in safety mechanisms. This prevents scenarios where a model appears
safe
because a BOS detector is triggered, but actually fails on adversarial paraphrases or rephrased harmful prompts. The system can also generate a report distinguishing positional vs. content features for human auditors.
Improvement 2: Feature Selection for Interpretability-Driven Steering
Improvement 3: Automated Feature Categorization for Mechanistic Interpretability
Improvement 4: Robustness Testing Against Token-Level Shortcuts
Improvement 5: Cross-Model Feature Transferability Assessment
Abstract
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce Feature Nonlocality (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in 73 -- 84% of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by 4.6 points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.
Sources
- Remarks on the disproof of the unit distance conjecture
- Sparse Autoencoders are Capable LLM Jailbreak Mitigators
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
- Training Verifiers to Solve Math Word Problems
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Scaling and evaluating sparse autoencoders
- A mathematical perspective on Transformers
- OpenThoughts: Data Recipes for Reasoning Models
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Measuring Mathematical Problem Solving With the MATH Dataset
- Rigorously Assessing Natural Language Explanations of Neurons
- The Stack: 3 TB of permissively licensed source code
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Let's Verify Step by Step
- Agentic Publication Protocol: An Attempt to Modernize Scientific Publication
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection