Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv:2608.10537 · cs.AI, cs.LG · Submitted 2026-08-11 · Read on arXiv

Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi

University of Oxford · Stanford University

cs.AI, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 18 pages, 9 figures

Code: https://github.com/lccqqqqq/sae-feature-nonlocality

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper addresses a central challenge in Mechanistic Interpretability (MI): distinguishing surface-level, token-driven features from genuinely high-level, abstract features in Sparse Autoencoders

Terminology

Summary

The paper addresses a central challenge in Mechanistic Interpretability (MI): distinguishing surface-level, token-driven features from genuinely high-level, abstract features in Sparse Autoencoders (SAEs). While SAEs have helped uncover mechanistic explanations for LLM behaviors such as reasoning and jailbreaking, existing feature selection methods—including keyword filtering, contrastive dataset curation, and auto-interpretation—do not adequately distinguish mechanistic complexity of features. The authors note that successful control of a target behaviour by intervening on a feature does not imply understanding of the underlying mechanism, citing examples where steering on a feature might simply boost token probabilities (e.g., Wait) rather than represent genuine uncertainty detection.

The paper introduces Feature Nonlocality (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. Formally:

For a feature a at layer l with activation z a(T, P) at position T of prompt P, the per-position influence is computed as the squared gradient norm:

J a(t, T, P) = ‖∂z a(T, P)/∂x t‖2

where x t are token representations at the layer-0 residual stream. This is normalized to a probability distribution over prefix positions, and the per-prompt nonlocality is the Shannon entropy:

H(a, P) = −Σ p(P) a*(t) log2 p(P) a*(t)

The dataset-level FNL averages this over firing events. The intuition is that Token-level features (e.g. bigram statistics) depend on local prefix cues and have low nonlocality, while abstract, topical features draw on contextual evidence spread broadly across the window and have high nonlocality. The definition draws inspiration from holographic duality in theoretical physics.

FNL rankings are stable across topically diverse corpora (WikiText, GSM8K, Code-Python). For Gemma-2-2B, Spearman correlations between corpora range from 0.710 to 0.918 across layers 5, 12, and 20. The supplementary material extends this to Llama-3-8B and Qwen3-8B, reporting disattenuated correlations (ρ̃) of 0.531–0.981, showing the ranking retains the majority of its attainable agreement at every model and depth.

FNL rises over the first half of the network and saturates at middle layers (l ≈ 13–16 for Gemma-2-2B), broadly following the corresponding residual-stream statistic. This pattern is robust across models and corpora.

FNL varies smoothly with decoder cosine similarity: features close together under decoder cosine similarity occupy similar ranges of nonlocality. The supplementary material shows this correlation is with write-side geometry rather than semantic similarity—semantic neighbors (by auto-interpretation embedding cosine) spread widely in FNL, while decoder neighbors cluster narrowly.

Using the methodology from Ma et al. (2026), the authors test 100 reasoning features in DeepSeek-R1-Distill-Llama-8B. FNL correlates negatively with activation recovery under token injection (Spearman ρ = −0.39 to −0.46 across layers 10–12 and 19), and AUC values of 0.73–0.84 for distinguishing token-driven from context-dependent features. The mean FNL for context-dependent features (4.57–4.71 bits) consistently exceeds that for token-driven features (4.03–4.38 bits).

FNL correlates positively with paraphrase robustness score S a (ρ = 0.27, p = 0.011, n = 96 features), where S a measures retention of activation under meaning-preserving paraphrases versus token shuffles.

Representative features show a clear continuum: single-token detectors (e.g., f608 Robert indicator, H(a) = 0.98 bits) sit at the low end, while epistemic hedging features (e.g., f13603, H(a) = 5.09 bits) sit at the high end. The paper concludes FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature.

The authors reproduce the CC-Delta pipeline (Assogba et al. 2026) for jailbreak-mitigation features in DeepSeek-R1-Distill-Llama-8B. Strikingly, 21 of the 25 selected features have 0 nonlocality—they activate almost exclusively at position-0 of plain harmful prompts, functioning as positional/BOS indicators rather than content recognizers. Steering experiments confirm: the positional subset (21 features) achieves held-out safety of 0.911 versus 0.504 for the content subset (4 features), with the unsteered baseline at 0.511. This reveals the features should be understood as surface-level, beginning-of-sequence (BOS) indicators rather than encoding the computation for recognizing harmful intents.

For DeepSeek-R1-Distill-Llama-8B at layer 19, steering the top 20% FNL features improves MATH-500 accuracy by 4.6 points over unsteered baseline (0.911 vs. 0.865), outperforming low-FNL steering (+3.8 points), random features (+3.6 points), and a representative ReasonScore-selected feature (+3.9 points). However, the authors caution this is a proof of concept rather than evidence that FNL reliably predicts steering utility—supplementary experiments show that outside DeepSeek-R1-Distill-Llama-8B, both high and low FNL arms fall below baselines.

  1. FNL is a stable, label-free, gradient-based metric requiring no contrastive datasets or interventions.

  2. FNL correlates with existing proxy measures of semantic abstractness (token-injection recovery and paraphrase robustness).

  3. FNL successfully distinguishes context-dependent reasoning features from token-driven ones in 73–84% of random pairs.

  4. FNL provides diagnostic value for safety-relevant behaviors, revealing that effective jailbreak-mitigation features are often positional rather than content-aware.

  5. FNL shows promise as a feature selection criterion for steering, though gains are model-specific.

The paper concludes that FNL helps distinguish what an intervention achieves and what the intervened feature represents—the former validated by steering utility, the latter requiring additional evidence that FNL supplies.

Improvements for AI systems

Improvement 1: Safety-Feature Auditing for Jailbreak Mitigation

  • What I can do: Automatically screen any SAE feature selected for safety interventions (e.g., refusal, harmlessness, or jailbreak-mitigation features) by computing its FNL score. If a feature has near-zero nonlocality (e.g., activates only at position 0 or on BOS tokens), I can flag it as a positional artifact rather than a content recognizer.

  • Improved system capability: An AI safety pipeline that, before deploying a steering vector or activation patch, rejects features with FNL below a threshold (e.g., < 0.5 bits) to avoid false confidence in safety mechanisms. This prevents scenarios where a model appears safe because a BOS detector is triggered, but actually fails on adversarial paraphrases or rephrased harmful prompts. The system can also generate a report distinguishing positional vs. content features for human auditors.

Improvement 2: Feature Selection for Interpretability-Driven Steering

Improvement 3: Automated Feature Categorization for Mechanistic Interpretability

Improvement 4: Robustness Testing Against Token-Level Shortcuts

Improvement 5: Cross-Model Feature Transferability Assessment

Abstract

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce Feature Nonlocality (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in 73 -- 84% of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by 4.6 points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.

Sources

Related papers