Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
cs.LG, cs.AI, cs.CL
Submitted: 2026-05-21
Updated: 2026-10-05
Comments: Focused analysis on attribution maps revealed different behaviors than the ones reported
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood.
Terminology
Abstract
Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAEs), which decompose static residual representations and miss the functional updates that drive cross-modal interaction. We adopt a function-centric framework based on Transcoders, sparse approximations of MLP sublayers that act as a causal proxy for layer-wise computation. Applied to Gemma 3-4B-IT, the framework decomposes the model into interpretable computational pathways linking image patches to directions in token generation. Transcoder attributions produce stronger and more stable effects on visually grounded tokens under patch ablation than SAE attributions, and align better with semantically relevant image regions. A False Visual Grounding counterfactual analysis confirms that the recovered pathways are specific to vision-language interaction.Finally, we perform a structural analysis of hallucinated generations, by extracting graph-based indicators from circuit traces produced by the transcoders. A logistic classifier over these mechanistic graph features predicts hallucinations at AUC 0.68. These results show that function-centric circuit decomposition yields interpretable and predictive accounts of multimodal computation in VLMs.
Sources
- Qwen3-VL Technical Report
- Mechanistic Interpretability for AI Safety -- A Review
- Uncertainty quantification for stationary and time-dependent PDEs subject to Gevrey regular random domain deformations
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
- Scaling and evaluating sparse autoencoders
- Gemma 3 Technical Report
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- SAE-V: Interpreting Multimodal Models for Enhanced Alignment
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- OpenAI GPT-5 System Card
- Interpretable Steering of Large Language Models with Feature Guided Activation Additions
- Interpretable and Testable Vision Features via Sparse Autoencoders
- Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
- Interpreting CLIP with Hierarchical Sparse Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks