Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
cs.LG, cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Explorations of Self-Repair in Language Models
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Transformer Circuit Faithfulness Metrics are not Robust
- How to use and interpret activation patching
- Transformer Feed-Forward Layers Are Key-Value Memories
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation
- Causal Abstractions of Neural Networks
- Representation Engineering: A Top-Down Approach to AI Transparency
- An Interpretability Illusion for BERT
- Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
- Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits
- Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning
- Through the Looking Glass: Directly Reading and Writing Transformers
- Optimal ablation for interpretability
- A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
- The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
- Learning Sparse Neural Networks through $L_0$ Regularization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks