VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
cs.CV, cs.AI
Submitted: 2025-07-04
Updated: 2025-11-28
Journal ref: Advances in Neural Information Processing Systems 38 (NeurIPS 2025)
DOI: 10.52202/085713-2642
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-Language Models (VLMs) excel at complex visual tasks such as VQA and chart understanding, yet recent work suggests they struggle with simple perceptual tests.
Terminology
Abstract
Vision-Language Models (VLMs) excel at complex visual tasks such as VQA and chart understanding, yet recent work suggests they struggle with simple perceptual tests. We present an evaluation of vision-language models' capacity for nonlocal visual reasoning: reasoning that requires chaining evidence collected from multiple, possibly distant regions of an image. We isolate three distinct forms of nonlocal vision: comparative perception, which demands holding two images in working memory and comparing them; saccadic search, which requires making discrete, evidence-driven jumps to locate successive targets; and smooth visual search, which involves following a continuous contour. Flagship models (e.g., GPT-5, Gemini 2.5 Pro, Claude Sonnet 4), even those that perform well on prior primitive-vision benchmarks, fail these tests and barely exceed random accuracy on two variants of our tasks that are trivial for humans. Our structured evaluation suite allows us to test whether VLMs can perform visual algorithms similar to those used by humans. Our findings show that despite gains in raw visual acuity, current models lack core visual reasoning capabilities.
Sources
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Vision language models are blind: Failing to translate detailed visual features into words
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2.5-VL Technical Report
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
- Slow Perception: Let's Perceive Geometric Figures Step-by-step
- VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
- MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models