Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs
cs.LG, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 26 pages, 12 figures, 3 tables. Code: https://github.com/edwardF8/Circuit-Diff | Data: https://huggingface.co/datasets/edwardF8/Circuit-Diff-paper-data | Demo: https://circuit-diff-paper.vercel.app/
Code: https://github.com/edwardF8/Circuit-Diff
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation.
Terminology
Abstract
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
Sources
- ADAG: Automatically Describing Attribution Graphs
- Linear Algebraic Structure of Word Senses, with Applications to Polysemy
- Automated Attribution Graph Interpretation via Probe Prompting
- Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Transcoders Find Interpretable LLM Feature Circuits
- Toy Models of Superposition
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models
- Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
- Rigorously Assessing Natural Language Explanations of Neurons
- CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
- DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
- Locating and Editing Factual Associations in GPT
- Mass-Editing Memory in a Transformer
- Understanding Factual Recall in Transformers via Associative Memories
- LLMs Can Annotate Attribution Graphs
- Automatically Interpreting Millions of Features in Large Language Models
- Long-form evaluation of model editing
- Gemma 2: Improving Open Language Models at a Practical Size
- Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks