How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "How Far Do Auto-Interpretation Labels Generalize".
Jane: Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature represents.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Now, let’s get into the main finding of this study, which is summarized in "How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings." Essentially, they asked if a feature labeled for a concept actually tracks that concept when you change the language or script.
Jane: What they found was that these auto-interpretation labels often fail to keep pace with the semantic generalization across different forms of content. Specifically, features whose labels describe semantic content miss the same meaning in Serbian up to four times more often than they do within English.
Lu: That four times multiplier is significant; it shows a systematic failure where the label doesn't capture the concept accurately in that specific language variant.
Meng: So, even when we know a feature is supposed to be tracking something meaningful, its interpretation gets weaker as we move away from the dominant training distribution.
Lalam: It’s important to understand this gap because it tells us that an auto-interpretation label doesn't automatically guarantee how a feature will behave on the same concept in any other form of expression.
Tom: And they also observed a trend related to the structure of the model itself, specifically that network depth plays a role in how much these features miss meaning.
Jane: They found that features miss Serbian Cyrillic up to one point thirty-five times more than Serbian Latin, even though those two scripts are essentially deterministic transliterations of each other.
Lu: That disparity between the two scripts, which are so closely related structurally, points toward a dependency on the specific script's representation in the training data rather than pure semantic structure.
Meng: It sounds like deeper layers become more specialized and less robust when it comes to cross-script interpretation of features.
Lalam: If that pattern holds as depth increases, it means we should be cautious when relying on deep layer labels for cross-lingual understanding.
The paper's summary: Tom: Moving past the findings, the authors suggest several ways to improve how we use these auto-interpretation labels and what they tell us about the model’s internal structure.
Jane: One key improvement suggested is to integrate a secondary classifier, similar to what they used, that can distinguish between a feature's label as a genuine content claim versus just a surface or language claim.
Lu: That’s essentially building a verification layer into the pipeline so we don't just accept labels blindly but check if they have semantic weight first.
Meng: From an implementation side, that means developing a mechanism that assigns a confidence score to every auto-interpretation label based on whether it passes that content claim filter.
Lalam: That sounds like it gives us a much more reliable interface for the model because we can trust the labels we see with higher certainty.
Tom: They also proposed developing a dynamic labeling protocol where features must pass a cross-lingual content check before they are assigned an auto-interpretation label in any scenario.
Jane: That would mean that if a feature is active in English but fails to fire when presented with equivalent Serbian content, we should flag that discrepancy for deeper investigation.
Lu: It pushes the research toward creating systems that actively probe for these cross-lingual inconsistencies rather than just reporting what the model *thinks* it’s doing.
Meng: That moves us closer to building agents that can actually detect when their interpretation of a concept breaks down across different input modalities or languages.
The paper's improvements: Tom: So, to wrap up, the main implication from this study on "How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings" is that while SAE features do encode abstract meaning that survives surface form changes, the auto-interpretation labels often overpromise.
Jane: They don't guarantee consistency when you switch contexts because the failure rates are significantly higher in specific language variants compared to others.
Lu: The authors conclude that these failures align with the estimated training coverage, which suggests the unreliability is graded and consistent with how much each form of text was present in the pretraining data.
Meng: So, practically speaking, this means we can’t take an auto-interpretation label as a universal guarantee about a concept's meaning across all languages and scripts.
Lalam: I think we need to treat these labels as claims about how a feature performs on the inputs it saw most often during training, not as absolute truths about the concept itself.
Tom: Exactly, so the message for listeners is that these labels are guides for investigation, not final answers when you’re dealing with multilingual models.
Jane: It’s a very nuanced finding because it shows where the model's understanding is strongest and where its interpretations are most fragile.
Lu: I think this opens up a lot of avenues for us to design more robust interpretability tools that account for these specific linguistic biases we just saw.
Meng: For practical application, this means we need layered systems that can assess the reliability of an interpretation based on how deep in the network the feature is firing.
Lalam: Understanding this graded pattern helps us build AI that is more aware of its own limitations when interpreting things across different cultural and linguistic contexts.
Conclusion: Tom: So, to wrap up on this study about "How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings," we’ve seen that while SAE features capture abstract meaning, the auto-interpretation labels they generate often fail to generalize consistently across different languages and scripts.
Jane: It really highlights a specific challenge in interpreting what these models are actually "thinking" versus what they are just reflecting from their training data distribution.
Lu: The authors showed that the failure rate for Serbian Cyrillic was much higher than for Latin, even though those scripts are closely related, which points to how much the model relies on the specific script it saw most often.
Meng: From a practical standpoint, this means if we’re deploying an AI system that relies on these labels for something critical, we need to know exactly where that reliability drops off based on what language it’s encountering.
Lalam: It gives me a vision where we can build cultural understanding tools that are more aware of how different scripts shape the representation of ideas in our society.
Tom: Right, and the core finding is that these labels don't match the actual semantic content reliably across all forms of expression.
Jane: It’s a crucial distinction between what a model *can* do and what it *reliably* does under pressure when inputs change.
Lu: The next step, which they suggest, is for us to develop dynamic protocols that require features to pass a content check before we even assign an interpretation label.
Meng: That would be smart; it forces the AI to prove its semantic grounding before we trust its surface-level labels for high-stakes decisions.
Lalam: If we can improve this, imagine AI agents that can navigate cultural nuances with far greater accuracy and respect.
Tom: Absolutely, that moves us from just looking at what the model says to understanding the actual structure of its knowledge.
Jane: This paper really makes you think about how much we trust the "interpretation" layer versus the raw feature activations underneath it.
Lu: I’m looking forward to seeing how this graded pattern of unreliability affects our future work on cross-lingual semantic tracking.
Meng: That graded nature is what keeps me interested; it shows us exactly where we need to focus our computational resources for better robustness.
Lalam: It’s exciting because it suggests that even in the most abstract layer, there are deep structures that just need a little bit of context to unlock their full potential.
Tom: Well said, Lalam, and let’s keep this momentum going as we look at how these insights apply to other challenging papers on arXiv next week.
Columbia University
cs.CL
Submitted: 2026-05-29
Updated: 2026-10-07
Code: https://github.com/Sripadkarne/auto-interp-cross-lingual-eval
Importance score: 92/100
The gist: Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature
Key concepts
- Sparse Autoencoder (SAE) Features
- These are mathematical representations learned by a model that can be interpreted using natural language. The paper uses these features to see if the underlying structure captures abstract meaning that is consistent across different languages and writing systems.
- Cross-Lingual Semantic Features
- This refers to whether a feature in the AI model encodes an idea or concept that remains meaningful even when expressed in different languages, scripts, or wordings. The study found evidence for these features existing at the feature level.
- Auto-Interpretation Labels
- These are the human-readable descriptions generated by the AI to explain what a specific SAE feature does. The key finding is that these labels are unreliable; they describe how a feature behaves on languages and scripts it has seen most often, not necessarily the true concept.
- Controlled Factorial Paradigm
- This was the experimental setup used to systematically isolate variables. Researchers compared content across different script variations (e.g., Latin vs. Cyrillic), languages, and paraphrases to determine which form of variation affects the feature's activation.
Terminology
Summary
Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature represents. The gist: Auto-interpretation labels reflect a feature’s behavior on the languages and scripts a model has seen most in training, rather than the concept itself.
Testing Cross-Lingual Semantic Features
The study investigates whether SAE features encode abstract meaning that generalizes across different languages, scripts, and wordings by using Serbian digraphia as a controlled testbed. The researchers first found that SAE feature sets activated by the same content in different languages, scripts, and wordings share substantial overlap (mean Jaccard 0.39 vs. 0.13 random baseline).
This suggests genuine cross-lingual semantic features
exist at the feature level, which is robust across model scales and architectures.
Evaluating Label Generalization
The core question addressed is whether auto-interpretation labels keep pace with this semantic generalization across different forms of the content. The findings indicate that "features whose labels describe semantic content miss the same meaning in Serbian up to 4× more often than within English, and miss Serbian Cyrillic more than Serbian Latin—two scripts that are deterministic transliterations of each other. The authors conclude that
auto-interpretation labels reflect a feature’s behavior on the languages and scripts a model has seen most in training, rather than the concept itself."
Methodology: Controlled Factorial Paradigm
The researchers constructed a factorial paradigm using 300 sentences from the FLORES+ dataset, which includes four language-script variants (English in Latin, Serbian in Cyrillic/Latin, Russian in Cyrillic) and three conditions (original translation, meaning-preserving paraphrase, and random partner). This setup isolates script, language, wording, and meaning independently. For example:
-
Script comparison:
Sr-Cyrillic orig vs. Sr-Latin orig (same language and content, only script varies).
-
Language comparison:
Sr-Cyrillic orig vs. Ru-Cyrillic orig (same script and content, language varies).
-
Meaning comparison:
English orig vs. Russian paraphrase (script, language, and wording all differ; only meaning shared).
Feature Selection and Content Claim Filtering
To test the labels, the study first selected features that passed a co-activation filter: they must fire on both the English original and the Russian paraphrase to show strong evidence the feature tracks meaning rather than any surface property.
Subsequently, a separate classifier LLM was used to classify each auto-interpretation label as a content claim (semantic topic or meaning) versus a surface or language claim,
discarding features whose labels describe only surface form.
Results and Implications
The analysis revealed that while SAE features appear to encode abstract meaning that largely survives a complete change of surface form,
the auto-interpretation labels often fail. The miss rates
for content-labeled features are significantly higher in Serbian Cyrillic compared to English or Russian, with the gap growing with network depth (e.g., reaching 1.35× at layer 53). This failure is consistent with training distribution, suggesting that the Serbian Cyrillic script is comparatively scarce in the web text that dominates modern pretraining corpora,
leading to a graded pattern of unreliability where reliability degrades across one as depth increases.
The practical lesson is that an auto-interpretation label guarantees nothing about how a feature behaves on the same concept in another form.
Robustness Checks
To ensure findings are not artifacts of specific extraction setups, the study replicated Leg 1 decomposition across different model scales (Gemma-3-1B, Gemma-3-12B) and architectures (Llama-3.1-8B). The qualitative structure replicates
across these variations, confirming that the factorial decomposition is robust to model scale and architecture. Furthermore, testing on different pooling strategies (last-token vs. mean vs. max pooling) showed that only last-token cleanly separates the conditions from their baselines,
supporting its use for isolating surface form effects. Finally, analysis of tokenization differences confirmed that tokenization asymmetry is not a meaningful driver of the cross-script patterns we report.
Conclusion
The study concludes that SAE features encode genuine semantic content, but these content labels overpromise by failing to generalize across languages and scripts. The failures are graded, aligning with the estimated representation of each form in training data, and tend to sharpen with network depth. Therefore, auto-interpretation labels should be understood as claims about a feature’s behavior on its best-represented inputs, not as guarantees about the concept in general.
Limitations
The study notes that while the graded miss-rate pattern is observed across three languages, it has not been verified for more typologically distant and lower-resource languages.
Improvements for AI systems
Here are the specific improvements to AI systems that can be derived from this research, categorized by the aspect of improvement:
) 1. Improved Feature Interpretation Reliability (The Core Finding)
By implementing a verification layer for auto-generated labels, AI systems can move beyond simply displaying features to what they might mean
and instead provide what they reliably mean.
-
Specific Improvement: Integrate a secondary, content-claim classifier (like the one tested in Section C) into the auto-interpretation pipeline. This classifier must be trained to distinguish between semantic concepts (content-claim) and superficial attributes (surface/language/script claims).
-
Improved AI Capability: The system can output features with a confidence score indicating whether their label is likely a genuine semantic concept or merely a surface artifact related to the specific language/script rendering. This drastically reduces
false positives
in safety monitoring or concept tracking.
) 2. Robust Cross-Lingual Concept Tracking (The Generalization Gap)
The paper demonstrates that labels fail systematically on less-represented languages and scripts, and this failure is invisible from the label itself.
-
Specific Improvement: Develop a dynamic labeling protocol that requires features to pass a
cross-lingual content check
(firing on content in both English and Russian) before they are assigned a potentially unreliable auto-interpretation label. -
Improved AI Capability: When monitoring for concepts like
deception,
the system can flag instances where the feature is active in English but fails to fire when presented with equivalent content in Serbian Cyrillic, even if the model's internal reasoning suggests it should. This allows practitioners to understand the concept's truesemantic range
across different forms of expression, rather than relying on a single language's performance.
) 3. Depth-Aware Feature Prioritization (The Scaling Effect)
The analysis shows that reliability changes with network depth (deeper layers become more specialized/less language-agnostic).
-
Specific Improvement: Implement a layer-specific reliability weighting function for auto-interpretation labels. Features from deeper layers in the network should receive lower weight when used for high-stakes concept tracking, unless the context strongly suggests a need to probe specific surface forms.
-
Improved AI Capability: For complex reasoning tasks, the system can prioritize features from earlier layers where semantic signals are more robust across linguistic variations, while using later layers only for fine-grained output generation or surface pattern matching.
) 4. Tokenization Asymmetry Mitigation (The Technical Finding)
The research confirms that tokenization differences (e.g., Cyrillic characters costing more tokens than Latin-derived Cyrillic) do not drive the cross-script feature overlap patterns, but they do affect raw token counts.
-
Specific Improvement: Implement a dynamic token normalization layer specifically for cross-script comparisons during feature similarity calculations. Instead of comparing raw feature Jaccard scores directly, normalize them against the expected subword token expansion ratio (e.g., the 15% overhead observed for Serbian Cyrillic).
-
Improved AI Capability: When comparing feature sets between languages written in different scripts (like Latin vs. Cyrillic), the system can be less susceptible to artifacts caused by tokenization differences, ensuring that observed semantic similarity is truly a reflection of cross-lingual meaning rather than an artifact of how the model tokenizes different character sets.
) 5. Adaptive Labeler Selection (The Practical Takeaway)
The comparison between two labelers (Claude Sonnet vs. Gemini Flash Lite) shows they converge on the same results, but their strictness differs by layer.
-
Specific Improvement: Create a meta-classifier that assesses the confidence/strictness of different auto-interpretation labelers based on layer depth and context.
-
Improved AI Capability: The system can automatically select the most reliable labeler for a given feature at a given network depth, ensuring that critical safety or auditing decisions are always based on the most conservative (strictest) interpretation available at that point in the model's processing.
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Scaling and evaluating sparse autoencoders
- Gemma 3 Technical Report
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Rigorously Assessing Natural Language Explanations of Neurons
- Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
- NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering