Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits
cs.LG, cs.AI
Submitted: 2026-07-02
Updated: 2026-09-23
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mechanistic interpretability seeks to explain transformer behavior through circuits: sets of internal components that causally support a behavior.
Terminology
Abstract
Mechanistic interpretability seeks to explain transformer behavior through circuits: sets of internal components that causally support a behavior. However, self-repair creates a blind spot: ablating a primary component can activate a dormant backup, so a circuit that explains behavior in the intact model can become incomplete under the intervention used to test it. We formulate this gap as conditional circuit completion: given a primary set, identify components that become causally important after its removal. We introduce conditional co-ablation (CoAx), which ranks candidates by growth in ablation effect after primary-set removal. We show that a perfectly dormant backup can be indistinguishable from an irrelevant component to per-unit intact-state scores, whereas its conditional effect change exactly aggregates all interaction orders linking it to the removed set. On GPT-2-small's Indirect Object Identification (IOI) circuit, CoAx recovers the documented backup heads at 0.941 ROC-AUC, versus 0.815 for the strongest intact-state attribution baseline and 0.758 for the matched conditional-energy control. Recovery drops to 0.40 +/- 0.13 AUC for alternative component sets matched in behavioral effect, output displacement, and depth, showing that recovery is specific to the removed circuit. Beyond recovery, the CoAx-selected heads are causally load-bearing: freezing them after primary removal sharply reduces the IOI margin, while adding them to the incomplete circuit reduces incompleteness from 0.75 to 0.21. More broadly, conditional growth aligns with intervention-derived repair in 11/12 held-out instances across 4 mechanism clusters, and CoAx completions outperform matched random completions on all 8 non-GPT-2 models spanning 6 architecture families. Together, causal explanations of self-repairing transformers must account for backup circuitry when primary components fail.
Sources
- Mechanistic Interpretability for AI Safety -- A Review
- Finding Transformer Circuits with Edge Pruning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- Gemma 2: Improving Open Language Models at a Practical Size
- Neuron Shapley: Discovering the Responsible Neurons
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
- Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
- A Fast Post-Training Pruning Framework for Transformers
- The Llama 3 Herd of Models
- Transformer Circuit Faithfulness Metrics are not Robust
- Importance Estimation for Neural Network Pruning
- 2 OLMo 2 Furious
- Qwen2.5 Technical Report
- Explorations of Self-Repair in Language Models
- Axiomatic Attribution for Deep Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks