Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Seeing Isn't Knowing".
Jane: Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we’ve talked about how these models get overconfident when things are occluded or perspectives are confusing in this paper called "Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?", and it really zeroes in on that assumption that visual observations are always sufficient.
Jane: They claim that because visual data is inherently limited—due to things like occlusion hiding objects or perspective making geometry misleading—current spatial reasoning benchmarks often fail to test if models can actually recognize when a question is unanswerable.
Lu: The central thesis of this work is that we need a controlled framework, SPATIALUNCERTAIN, to see if models can identify when they should abstain from answering instead of guessing when the visual evidence is incomplete or misleading.
Meng: It’s about moving the focus from just answer correctness to recognizing uncertainty and figuring out what extra observations are necessary before making a call.
Lalam: The paper sets up two main challenges: occlusion, which hides target information, and perspective ambiguity, which introduces misleading visual cues that don't change the underlying geometry.
Tom: Under these conditions, they design spatial questions—visibility, relative position, depth ordering—that are answerable in a clean view but become unanswerable when those specific observational challenges are introduced.
Jane: The paper shows that under occlusion or perspective ambiguity, questions requiring access to hidden targets or relying on visual appearance alone become impossible to answer reliably.
Lu: They set up evaluation tasks like ViewSel and AbstainViewSel, which test the model's ability to select an informative viewpoint when things are ambiguous and whether it can recognize unreliability in the first place.
Meng: The framework is designed to systematically manipulate these observational conditions so we can pinpoint exactly where models start failing to handle spatial reasoning correctly.
Lalam: The main point they make is that models struggle not only to abstain but also to identify which alternative viewpoints would provide reliable evidence when visual cues become misleading under perspective ambiguity.
Tom: This points out a real limitation in current systems: they can’t tell the difference between bad data and good data, especially when the visual information itself starts giving them false hints.
Jane: It matters because it challenges the way we've been testing these models—if we don't test for uncertainty, we don're not truly understanding their spatial capabilities in messy environments.
Lu: This work provides a concrete way to challenge that assumption by building a system where the model’s required behavior is to recognize when it cannot determine the truth.
Meng: It helps ground the discussion because it shows exactly where the current systems break down when they move from perfect, clean data to real-world, imperfect data.
Lalam: Ultimately, this research is about establishing a new standard for how we evaluate spatial reasoning by demanding that models demonstrate awareness of their own observational limits.
Conclusion: Tom: So wrapping up this discussion on "Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?", the authors are Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor, and Mohit Bansal. They really challenge us to think about what it means for an AI to be smart in a physical world.
Jane: The paper suggests that the future of spatial AI isn't just about increasing raw visual data; it’s more about teaching models how to manage uncertainty and know when to pause their reasoning process.
Lu: If these models can learn to recognize when their visual input is unreliable, we open up possibilities for more robust AI agents that operate in unpredictable, real-world settings where perfect conditions aren't the norm.
Meng: For practical applications, this means we could build systems that are far safer because they won't confidently make decisions based on shaky visual evidence when they should instead ask for clarification or stop.
Lalam: The impact is that we’re moving toward a more mature understanding of spatial reasoning, where models don't just output numbers but understand the context and limitations of their perception.
Tom: It really shifts the focus from achieving perfect accuracy under ideal conditions to building systems that can handle the messy reality of visual data, which is where most real-world AI lives.
Jane: So, we’re looking at a future where spatial intelligence involves not just seeing what’s there, but intelligently assessing whether what they see is trustworthy enough to be used for a decision.
Lu: This opens up avenues for creative applications in areas that rely on navigation or manipulation where failure due to overconfidence could have real consequences.
Meng: It gives us a clear direction on how to fine-tune these models—we need diversity in their training data, especially around visual ambiguity, to make them better at this critical skill.
Lalam: If we can nail this ability to abstain and seek reliable evidence, it could lead to AI that is far more reliable for complex tasks that require real-time decision-making.
UNC Chapel Hill · Google Research
cs.CV, cs.AI, cs.CL
Submitted: 2026-05-28
Updated: 2026-10-06
Comments: Website: https://zhangyuejoslin.github.io/spatialuncertain/
Project page: https://spatialuncertain.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 76/100
The gist: Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable, which this
Key concepts
- SPATIALUNCERTAIN
- A controlled evaluation framework built using 3D simulated environments. It systematically tests if a model can identify when visual observations are incomplete or misleading by manipulating scenes with occlusion and perspective shifts.
- Occlusion
- This simulates hiding parts of an object or scene from the camera's view. In this test, the occluder is placed along the line of sight, creating 'missing information' that challenges a model's ability to answer questions about hidden objects.
- Perspective Ambiguity
- This involves creating misleading visual cues by changing the viewpoint without changing the underlying geometry. For example, viewing an object from a slightly different angle can create systematic appearance differences that confuse models about its true size or shape.
Terminology
Summary
Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable, which this work challenges by constructing a controlled evaluation framework to test when models should abstain from answering spatial questions.
The gist: Models are prone to overconfident answering, attempting to solve spatial reasoning tasks even when visual evidence is incomplete or misleading, with average accuracy around 30% under occlusion and below 10% under perspective ambiguity.
SPATIALUNCERTAIN: Controlled Evaluation Framework
The paper introduces SPATIALUNCERTAIN, a controlled evaluation framework based on 3D simulated environments designed to evaluate whether models can recognize when visual observations are unreliable and identify additional informative evidence. This framework moves beyond assessing answer correctness by focusing on recognizing uncertainty and seeking reliable evidence.
The framework is constructed using diverse indoor scenes generated by Holodeck, allowing for systematic manipulation of observational conditions through two controlled perturbations:
-
Occlusion: This simulates hiding target information, leading to
missing information.
The pipeline involves selecting a target-occluder pair and placing the occluder along the line of sight, resulting inpartial or full occlusion
configurations. -
Perspective Ambiguity: This introduces misleading visual cues due to viewpoint bias. For object pairs, this involves generating a reference view (equidistant) and an ambiguous view (laterally shifted), which induces
systematic appearance differences without altering the underlying geometry.
Task Design and Answerability
The framework designs spatial questions whose answerability varies systematically with the observation conditions. The four question types considered are: Visibility, Relative position, Depth ordering, and Size/Shape. Under clean observations, all questions are answerable. However, under full occlusion or ambiguous perspectives, answerability changes: for example, under full occlusion, questions requiring access to the hidden target (relative position, depth, and size/shape) become unanswerable.
Similarly, under perspective ambiguity at an ambiguous view, questions about size and shape cannot be reliably answered from visual appearance alone.
Evaluation Tasks
To comprehensively assess model behavior under these challenges, two complementary evaluation tasks are introduced:
-
ViewSel (Viewpoint Selection): This single-stage task measures the ability to select an informative viewpoint when ambiguity exists. Models are presented with five candidate views and asked to identify the view that
best supports answering a spatial reasoning question about physical size.
-
AbstainViewSel (Two-Stage): This joint task evaluates both recognition of unreliability and selection of evidence. Stage 1 requires the model to answer based only on the biased view, including an option to abstain with
Cannot determine.
Stage 2 is triggered only if the model abstains, and the model must then select a reliable alternative viewpoint.
Findings on Model Failures
The evaluation across eight vision-language models reveals two major limitations:
-
Models are
prone to overconfident answering, attempting to solve spatial reasoning tasks even when visual evidence is incomplete or misleading,
with average accuracy around 30% under occlusion and below 10% under perspective ambiguity. -
Models struggle to identify which additional viewpoints would provide reliable evidence, as performance on AbstainViewSel drops sharply compared to ViewSel (e.g., GPT-5.4 decreases from 70.9 to 22.6).
Mitigation Strategies
The paper investigates two mitigation approaches: prompting and fine-tuning.
1 Structured Prompting: This guides the model to first assess object visibility and viewpoint reliability before answering, but it introduces a trade-off with answerable accuracy.
For instance, GPT-5-mini shows a substantial gain on Occ-Unans (7.8→30.4) at the cost of answerable accuracy (64.7→54.7).
2 Fine-tuning: The study finds that abstention is learnable but requires diversity.
Fine-tuning on diverse forms of visual ambiguity
using LoRA on both occlusion and perspective data substantially improves both answerable and unanswerable performance across conditions, resolving the trade-off observed with prompting alone. Single-condition training fails to generalize across ambiguity types.
Asymmetry of Visual Input
A key asymmetry is uncovered regarding the effect of visual input: visual information is beneficial when evidence is missing, improving both answering and abstention under occlusion, but can actively mislead models under perspective ambiguity.
Under perspective ambiguity, adding visual input often degrades models’ ability to recognize unanswerable cases,
suggesting current models struggle to assess the reliability of visual evidence.
Broader Impact
The findings suggest that current VLMs lack a unified understanding of observational reliability in spatial reasoning.
The work calls for moving beyond answer correctness toward evaluating whether models know when to abstain and how to seek reliable evidence, which has implications for reliability-critical applications like embodied agents.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the findings of SPATIALUNCERTAIN (Zhang et al., 2026) and identified several concrete, high-impact improvements for current Vision-Language Models (VLMs).
Here are the specific improvements and what these improved AI systems can achieve:
)1. Implement a Unified Observational Reliability
Module
The core failure mode is that VLMs lack a unified understanding of when visual evidence is reliable versus misleading.
-
Instead of just outputting an answer, the model should first pass its observation through a module that assesses the input quality based on learned cues for occlusion and perspective distortion.
-
This module should explicitly predict: (a) If the target is fully occluded/unseen, or (b) If geometric cues are unreliable due to viewpoint bias.
-
By integrating this reliability assessment into the reasoning pipeline, the system moves from
answering blindly
toreasoning based on evidence quality.
)2. Develop a Two-Stage Uncertainty Protocol (Abstain then Seek Evidence)
The current models fail not just by answering incorrectly but by failing to recognize when they shouldn't answer and failing to look for better views.
-
Implement the two-stage evaluation protocol: If Stage 1 (biased view) results in
Cannot determine,
trigger Stage 2, where the model is explicitly asked to select an informative viewpoint from a set of candidates. -
This creates a robust decision pathway: if evidence is bad, don't guess; instead, actively seek corrective data.
-
The resulting AI system can perform complex spatial tasks by intelligently requesting (or simulating) necessary camera moves or sensor reorientations to resolve ambiguity.
)3. Enhance Visual Input Sensitivity to Ambiguity (Asymmetric Calibration)
The paper reveals a critical asymmetry: visual input helps when information is missing (occlusion), but actively degrades performance when cues are misleading (perspective ambiguity).
-
Design a calibration mechanism that dynamically adjusts the weight of visual features based on the perceived reliability of the viewpoint. When perspective ambiguity is detected, this system should suppress the influence of potentially misleading appearance features and rely more heavily on viewpoint-invariant geometric constraints (like relative position or visibility).
-
This allows for
smart
visual processing—using sight when it's clear, but ignoring misleading sight when it's distorted.
)4. Move Beyond Simple Prompting to Diverse, Condition-Specific Fine-Tuning
The limitations of structured prompting are that they introduce an answer/abstention trade-off that is model-dependent and not robust across different ambiguity types.
-
Implement a fine-tuning strategy using LoRA (Low-Rank Adaptation) on a diverse dataset containing both occlusion and perspective ambiguity scenarios (LoRA-Mixed approach).
-
This training regimen teaches the model to generalize the concept of
unreliable observation
rather than just memorizing how to follow a prompt. -
The resulting system will possess a generalizable
abstention capability,
allowing it to reliably recognize when its input is fundamentally flawed, regardless of whether the flaw is occlusion or perspective distortion.
)What the Improved AI System Can Do (Specific Capabilities):
-
Organized Robotic Navigation: An embodied agent can navigate a complex environment (e.g., a kitchen or warehouse) by not just looking at objects, but by assessing if its current view is reliable enough to determine object depth or size accurately. If perspective ambiguity arises from a tight corner, the system will proactively request a wider sweep of vision before attempting to grasp an item, preventing collisions and dropped objects.
-
Advanced Inspection and Maintenance: A robotic inspection system can reliably identify subtle geometric differences (like the exact size comparison of two wall-mounted items) even when viewed from an oblique angle. It won't just give a potentially wrong answer; it will flag the specific viewpoint needed to confirm that measurement, ensuring high-precision quality control.
-
Safety and Risk Assessment: In a dynamic environment, if sensor data is partially occluded (e.g., a safety barrier), the system can correctly abstain from making an immediate movement decision (like driving through or opening a door) until sufficient, reliable visual evidence is gathered from a known good viewpoint. This prevents dangerous overconfident actions based on incomplete data.
-
Robust Human-Computer Interaction: In virtual reality interfaces, the AI can present information that is contextually honest. If a user asks
Is this shelf taller than that one?
, and the current camera angle makes the comparison ambiguous, the system will respond with a calibrated uncertainty statement (Cannot determine reliably from this view
) rather than guessing incorrectly.
Sources
- Qwen2.5-VL Technical Report
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- LoRA: Low-Rank Adaptation of Large Language Models
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- Language Models (Mostly) Know What They Know
- AI2-THOR: An Interactive 3D Environment for Visual AI
- SQA3D: Situated Question Answering in 3D Scenes
- GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
- OpenAI GPT-5 System Card
- Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
- Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
- SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models