ObserverBench: Testing Mechanistic Estimates for Intervention and Control
cs.LG, cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 28 pages, 4 figures. Code, benchmark, leaderboards, and submission interface: https://kwisatzh.github.io/observerbench/. Frozen artifact release: https://doi.org/10.5281/zenodo.22136091
Project page: https://kwisatzh.github.io/observerbench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring.
Terminology
Abstract
Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action. Theory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.
Sources
- The Bayesian Geometry of Transformer Attention
- Steering Language Models With Activation Engineering
- Representation Engineering: A Top-Down Approach to AI Transparency
- Activation Steering with a Feedback Controller
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
- pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
- The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
- Interpretability Can Be Actionable
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Copy Suppression: Comprehensively Understanding an Attention Head
- Explorations of Self-Repair in Language Models
- AI Control: Improving Safety Despite Intentional Subversion
- In-context Learning and Induction Heads
- Qwen2.5 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks