Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally
Terminology
Abstract
Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally drive model behavior. We demonstrate that this gap cannot be closed from the geometry of probe weights alone: the features geometrically aligned with probe direction need not be the ones the model uses, so causal relevance requires intervention. We introduce a feature-level diagnostic that decomposes a deployed True/False probe into sparse-autoencoder (SAE) features, ranks those features by both probe alignment and by gradient sensitivity of the model's behavior, and ablates the resulting shared, probe-only, and random feature sets under a coherence gate. On the truth probe of Buerger et al. (2024) (TTPD), applied in the instructed truth/deception setting of Long et al. (2025) for Gemma2-9B-Instruct, the two rankings overlap only weakly (about 12%, Spearman rho = 0.10), and ablation dissociates them sharply: features the probe shares with the model flip the output far more (up to 27%) than equally sized probe-only (6%) or random (1%) features at full coherence, while probe-only features instead perturb the probe's own readout. The dissociation holds across five seeds and a held-out split, and an activation-aware selection of features flips behavior nearly three times as often as the probe's geometric top features (17.6% vs. 6.1%). In this setting, therefore, the geometric projection of a probe's weight vector alone does not identify the features the model causally uses; however, combining probe information with feature activation statistics recovers substantially more behaviorally causal features, and coherence-gated SAE intervention is needed to separate them from probe readouts.
Sources
- Understanding intermediate layers using linear classifier probes
- Truth is Universal: Robust Detection of Lies in LLMs
- Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Interpreting Attention Layer Outputs with Sparse Autoencoders
- Optimal ablation for interpretability
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?
- Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States
- Do Input Gradients Highlight Discriminative Features?
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
- Gemma 2: Improving Open Language Models at a Practical Size
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection