Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?
cs.AI
Submitted: 2026-06-23
Updated: 2026-09-02
Code: https://github.com/Ziyu-Yao-NLP-Lab/LLM-Circuit-Explainer
Terminology
Sources
- The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
- Mechanistic Interpretability for AI Safety -- A Review
- A Primer on the Inner Workings of Transformer-based Language Models
- InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
- SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
- Language Models Use Trigonometry to Do Addition
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
- Tracr: Compiled Transformers as a Laboratory for Interpretability
- NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
- TracrBench: Generating Interpretability Testbeds with Large Language Models
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Thinking Like Transformers
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Automated Interpretability and Feature Discovery in Language Models with Agents
- MIB: A Mechanistic Interpretability Benchmark
- LLM Evaluators Recognize and Favor Their Own Generations
- Automatically Interpreting Millions of Features in Large Language Models
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection