The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence
cs.CV, cs.CL, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 17 pages (7-page main text plus technical appendix), 10 figures, 6 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual input.
Terminology
Abstract
Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual input. We show this assumption fails on 40%--97% of samples across six VLMs and three perceptual benchmarks: blurring the question-relevant visual region leaves the next-token distribution nearly unchanged. We name this phenomenon the Visual Insensitivity Gap and quantify it with a per-sample Visual Sensitivity Index (VSI). The gap is a property of samples, not of models: VSI ranks correlate across models (grand-mean Spearman rho=+0.40, permutation p<10-3), so the same samples are flagged insensitive by VLMs sharing no architectural detail beyond a contrastively pretrained vision tower. The mechanism is concrete: on the insensitive samples, a linear probe on each model's own vision tower distinguishes perturbed from clean images at 0.72--0.79 accuracy, yet the model's argmax token changes on only 2%--11% of the same samples, an encoder--LLM gap above 0.65 on every model. Mapping VSI's diagnostic utility cell by cell surfaces a strong regime (multi-choice reasoning on capable VLMs: AUROC=0.85--0.87) and a weak regime (well-calibrated factuality, where softmax confidence already leads). VSI is not a universal best abstention signal; it is a sample-intrinsic indicator of vision-ignoring failure, best used as a conditional ensemble component.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Hallucination of Multimodal Large Language Models: A Survey
- Building and better understanding vision-language models: insights and future directions
- A Survey on Hallucination in Large Vision-Language Models
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- Don't Blink: Evidence Collapse during Multimodal Reasoning
- Vision Language Models are Biased
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models