Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery
cs.CL
Submitted: 2026-06-04
Updated: 2026-08-26
Code: https://github.com/UKPLab/arxiv2026-phantom-specialization
Terminology
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- An Interpretability Illusion for BERT
- Finding Interpretable Prompt-Specific Circuits in Language Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Localizing Model Behavior with Path Patching
- How to use and interpret activation patching
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
- Distributed Specialization: Rare-Token Neurons in Large Language Models
- Repetitions are not all alike: distinct mechanisms sustain repetition in language models
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
- Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
- The Fair Language Model Paradox
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
- Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
- Gemma: Open Models Based on Gemini Research and Technology
- Gemma 2: Improving Open Language Models at a Practical Size
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Adversarial Circuit Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering