The Misery of Mechanistic Interpretability: A Formal Perspective
cs.LG, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Keywords: artificial intelligence; safety; interpretability; large language models; formal methods; verification; certification; robustness; sparse autoencoder; transcoder
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes.
Terminology
Abstract
Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes. To gain insights into these models, interpretable replacement networks (IRNs) are trained at all layers, exposing interpretable features through sparsely activated neurons. However, the faithfulness of an IRN is usually evaluated only empirically on clean data, and we show that even semantically minor input perturbations flip the dominant IRN features-and thus the human-understandable interpretation-across five open-weight model families (GPT-2 small, Gemma 2 2B, Gemma 3 1B, Llama 3.2 1B, R1-Distill-Qwen 1.5B). We propose the first formal verification framework for the faithfulness of an IRN, where reachability analysis certifies a sound upper bound of the faithfulness gap in adversarial scenarios. Moreover, we show that verification-aware training of IRNs substantially tightens this certified bound, restoring a feature-level interpretation that safety auditors can act on. Together, these results give, to the best of our knowledge, the first formal guarantees for mechanistic interpretability of large language models.
Sources
- Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
- Weight-sparse transformers have interpretable circuits
- Gemma 2: Improving Open Language Models at a Practical Size
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- The 6th International Verification of Neural Networks Competition (VNN-COMP 2025): Summary and Results
- Qwen2.5 Technical Report
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks