Causal Probing for Internal Visual Representations in Multimodal Large Language Models
cs.AI
Submitted: 2026-05-07
Updated: 2026-09-03
Comments: Accepted at EMNLP 2026 Main
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen2.5-VL Technical Report
- Representation Engineering for Large-Language Models: Survey and Research Challenges
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Understanding intermediate layers using linear classifier probes
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- The Internal State of an LLM Knows When It's Lying
- Activation Steering for Chain-of-Thought Compression
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
- Mass-Editing Memory in a Transformer
- The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Programming Refusal with Conditional Activation Steering
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Towards Understanding Sycophancy in Language Models
- Reducing Hallucinations in Vision-Language Models via Latent Space Steering
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection