Can Interpretation Predict Behavior on Unseen Data?
cs.LG, cs.AI, cs.CL
Submitted: 2025-07-08
Updated: 2026-09-12
Code: https://github.com/vli31/id-predict-ood
License: http://creativecommons.org/licenses/by/4.0/
The gist: Interpretability research often predicts model responses to targeted mechanistic interventions.
Terminology
Abstract
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Transformers on simple synthetic tasks, where perfect in-distribution accuracy is compatible with multiple OOD generalization rules. We successfully use attention patterns -- observed only on in-distribution data -- to predict which rule each model follows on OOD data. Our experiments decouple the mechanistic faithfulness of our interpretation from its predictive value; ablations reveal such internal patterns can suppress rather than support the rule they predict, showing observational analysis can forecast behavior even when causal analysis fails to support a simple cause-effect link. Our findings are a proof-of-concept for a new interpretability objective: understanding model internals to predict behavior and assess reliability under distribution shift.
Sources
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Transferring Inductive Biases through Knowledge Distillation
- On the Robustness of Interpretability Methods
- When Models Manipulate Manifolds: The Geometry of a Counting Task
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
- Sparse Autoencoders Trained on the Same Data Learn Different Features
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
- Memory-Augmented Recurrent Neural Networks Can Learn Generalized Dyck Languages
- Analyzing the Generalization and Reliability of Steering Vectors
- The Importance of Being Recurrent for Modeling Hierarchical Structure
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks