Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

arXiv:2605.30557 · cs.CV, cs.AI, cs.CL · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Seeing Isn't Knowing".

Jane: Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve talked about how these models get overconfident when things are occluded or perspectives are confusing in this paper called "Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?", and it really zeroes in on that assumption that visual observations are always sufficient.

Jane: They claim that because visual data is inherently limited—due to things like occlusion hiding objects or perspective making geometry misleading—current spatial reasoning benchmarks often fail to test if models can actually recognize when a question is unanswerable.

Lu: The central thesis of this work is that we need a controlled framework, SPATIALUNCERTAIN, to see if models can identify when they should abstain from answering instead of guessing when the visual evidence is incomplete or misleading.

Meng: It’s about moving the focus from just answer correctness to recognizing uncertainty and figuring out what extra observations are necessary before making a call.

Lalam: The paper sets up two main challenges: occlusion, which hides target information, and perspective ambiguity, which introduces misleading visual cues that don't change the underlying geometry.

Tom: Under these conditions, they design spatial questions—visibility, relative position, depth ordering—that are answerable in a clean view but become unanswerable when those specific observational challenges are introduced.

Jane: The paper shows that under occlusion or perspective ambiguity, questions requiring access to hidden targets or relying on visual appearance alone become impossible to answer reliably.

Lu: They set up evaluation tasks like ViewSel and AbstainViewSel, which test the model's ability to select an informative viewpoint when things are ambiguous and whether it can recognize unreliability in the first place.

Meng: The framework is designed to systematically manipulate these observational conditions so we can pinpoint exactly where models start failing to handle spatial reasoning correctly.

Lalam: The main point they make is that models struggle not only to abstain but also to identify which alternative viewpoints would provide reliable evidence when visual cues become misleading under perspective ambiguity.

Tom: This points out a real limitation in current systems: they can’t tell the difference between bad data and good data, especially when the visual information itself starts giving them false hints.

Jane: It matters because it challenges the way we've been testing these models—if we don't test for uncertainty, we don're not truly understanding their spatial capabilities in messy environments.

Lu: This work provides a concrete way to challenge that assumption by building a system where the model’s required behavior is to recognize when it cannot determine the truth.

Meng: It helps ground the discussion because it shows exactly where the current systems break down when they move from perfect, clean data to real-world, imperfect data.

Lalam: Ultimately, this research is about establishing a new standard for how we evaluate spatial reasoning by demanding that models demonstrate awareness of their own observational limits.

Conclusion: Tom: So wrapping up this discussion on "Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?", the authors are Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor, and Mohit Bansal. They really challenge us to think about what it means for an AI to be smart in a physical world.

Jane: The paper suggests that the future of spatial AI isn't just about increasing raw visual data; it’s more about teaching models how to manage uncertainty and know when to pause their reasoning process.

Lu: If these models can learn to recognize when their visual input is unreliable, we open up possibilities for more robust AI agents that operate in unpredictable, real-world settings where perfect conditions aren't the norm.

Meng: For practical applications, this means we could build systems that are far safer because they won't confidently make decisions based on shaky visual evidence when they should instead ask for clarification or stop.

Lalam: The impact is that we’re moving toward a more mature understanding of spatial reasoning, where models don't just output numbers but understand the context and limitations of their perception.

Tom: It really shifts the focus from achieving perfect accuracy under ideal conditions to building systems that can handle the messy reality of visual data, which is where most real-world AI lives.

Jane: So, we’re looking at a future where spatial intelligence involves not just seeing what’s there, but intelligently assessing whether what they see is trustworthy enough to be used for a decision.

Lu: This opens up avenues for creative applications in areas that rely on navigation or manipulation where failure due to overconfidence could have real consequences.

Meng: It gives us a clear direction on how to fine-tune these models—we need diversity in their training data, especially around visual ambiguity, to make them better at this critical skill.

Lalam: If we can nail this ability to abstain and seek reliable evidence, it could lead to AI that is far more reliable for complex tasks that require real-time decision-making.

UNC Chapel Hill · Google Research

cs.CV, cs.AI, cs.CL

Submitted: 2026-05-28

Updated: 2026-10-06

Comments: Website: https://zhangyuejoslin.github.io/spatialuncertain/

Project page: https://spatialuncertain.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 76/100

The gist: Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable, which this

Key concepts

SPATIALUNCERTAIN
A controlled evaluation framework built using 3D simulated environments. It systematically tests if a model can identify when visual observations are incomplete or misleading by manipulating scenes with occlusion and perspective shifts.
Occlusion
This simulates hiding parts of an object or scene from the camera's view. In this test, the occluder is placed along the line of sight, creating 'missing information' that challenges a model's ability to answer questions about hidden objects.
Perspective Ambiguity
This involves creating misleading visual cues by changing the viewpoint without changing the underlying geometry. For example, viewing an object from a slightly different angle can create systematic appearance differences that confuse models about its true size or shape.

Terminology

Summary

Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable, which this work challenges by constructing a controlled evaluation framework to test when models should abstain from answering spatial questions.

The gist: Models are prone to overconfident answering, attempting to solve spatial reasoning tasks even when visual evidence is incomplete or misleading, with average accuracy around 30% under occlusion and below 10% under perspective ambiguity.

SPATIALUNCERTAIN: Controlled Evaluation Framework

The paper introduces SPATIALUNCERTAIN, a controlled evaluation framework based on 3D simulated environments designed to evaluate whether models can recognize when visual observations are unreliable and identify additional informative evidence. This framework moves beyond assessing answer correctness by focusing on recognizing uncertainty and seeking reliable evidence.

The framework is constructed using diverse indoor scenes generated by Holodeck, allowing for systematic manipulation of observational conditions through two controlled perturbations:

  1. Occlusion: This simulates hiding target information, leading to missing information. The pipeline involves selecting a target-occluder pair and placing the occluder along the line of sight, resulting in partial or full occlusion configurations.

  2. Perspective Ambiguity: This introduces misleading visual cues due to viewpoint bias. For object pairs, this involves generating a reference view (equidistant) and an ambiguous view (laterally shifted), which induces systematic appearance differences without altering the underlying geometry.

Task Design and Answerability

The framework designs spatial questions whose answerability varies systematically with the observation conditions. The four question types considered are: Visibility, Relative position, Depth ordering, and Size/Shape. Under clean observations, all questions are answerable. However, under full occlusion or ambiguous perspectives, answerability changes: for example, under full occlusion, questions requiring access to the hidden target (relative position, depth, and size/shape) become unanswerable. Similarly, under perspective ambiguity at an ambiguous view, questions about size and shape cannot be reliably answered from visual appearance alone.

Evaluation Tasks

To comprehensively assess model behavior under these challenges, two complementary evaluation tasks are introduced:

  1. ViewSel (Viewpoint Selection): This single-stage task measures the ability to select an informative viewpoint when ambiguity exists. Models are presented with five candidate views and asked to identify the view that best supports answering a spatial reasoning question about physical size.

  2. AbstainViewSel (Two-Stage): This joint task evaluates both recognition of unreliability and selection of evidence. Stage 1 requires the model to answer based only on the biased view, including an option to abstain with Cannot determine. Stage 2 is triggered only if the model abstains, and the model must then select a reliable alternative viewpoint.

Findings on Model Failures

The evaluation across eight vision-language models reveals two major limitations:

  1. Models are prone to overconfident answering, attempting to solve spatial reasoning tasks even when visual evidence is incomplete or misleading, with average accuracy around 30% under occlusion and below 10% under perspective ambiguity.

  2. Models struggle to identify which additional viewpoints would provide reliable evidence, as performance on AbstainViewSel drops sharply compared to ViewSel (e.g., GPT-5.4 decreases from 70.9 to 22.6).

Mitigation Strategies

The paper investigates two mitigation approaches: prompting and fine-tuning.

1 Structured Prompting: This guides the model to first assess object visibility and viewpoint reliability before answering, but it introduces a trade-off with answerable accuracy. For instance, GPT-5-mini shows a substantial gain on Occ-Unans (7.8→30.4) at the cost of answerable accuracy (64.7→54.7).

2 Fine-tuning: The study finds that abstention is learnable but requires diversity. Fine-tuning on diverse forms of visual ambiguity using LoRA on both occlusion and perspective data substantially improves both answerable and unanswerable performance across conditions, resolving the trade-off observed with prompting alone. Single-condition training fails to generalize across ambiguity types.

Asymmetry of Visual Input

A key asymmetry is uncovered regarding the effect of visual input: visual information is beneficial when evidence is missing, improving both answering and abstention under occlusion, but can actively mislead models under perspective ambiguity. Under perspective ambiguity, adding visual input often degrades models’ ability to recognize unanswerable cases, suggesting current models struggle to assess the reliability of visual evidence.

Broader Impact

The findings suggest that current VLMs lack a unified understanding of observational reliability in spatial reasoning. The work calls for moving beyond answer correctness toward evaluating whether models know when to abstain and how to seek reliable evidence, which has implications for reliability-critical applications like embodied agents.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the findings of SPATIALUNCERTAIN (Zhang et al., 2026) and identified several concrete, high-impact improvements for current Vision-Language Models (VLMs).

Here are the specific improvements and what these improved AI systems can achieve:


)1. Implement a Unified Observational Reliability Module

The core failure mode is that VLMs lack a unified understanding of when visual evidence is reliable versus misleading.

  • Instead of just outputting an answer, the model should first pass its observation through a module that assesses the input quality based on learned cues for occlusion and perspective distortion.

  • This module should explicitly predict: (a) If the target is fully occluded/unseen, or (b) If geometric cues are unreliable due to viewpoint bias.

  • By integrating this reliability assessment into the reasoning pipeline, the system moves from answering blindly to reasoning based on evidence quality.

)2. Develop a Two-Stage Uncertainty Protocol (Abstain then Seek Evidence)

The current models fail not just by answering incorrectly but by failing to recognize when they shouldn't answer and failing to look for better views.

  • Implement the two-stage evaluation protocol: If Stage 1 (biased view) results in Cannot determine, trigger Stage 2, where the model is explicitly asked to select an informative viewpoint from a set of candidates.

  • This creates a robust decision pathway: if evidence is bad, don't guess; instead, actively seek corrective data.

  • The resulting AI system can perform complex spatial tasks by intelligently requesting (or simulating) necessary camera moves or sensor reorientations to resolve ambiguity.

)3. Enhance Visual Input Sensitivity to Ambiguity (Asymmetric Calibration)

The paper reveals a critical asymmetry: visual input helps when information is missing (occlusion), but actively degrades performance when cues are misleading (perspective ambiguity).

  • Design a calibration mechanism that dynamically adjusts the weight of visual features based on the perceived reliability of the viewpoint. When perspective ambiguity is detected, this system should suppress the influence of potentially misleading appearance features and rely more heavily on viewpoint-invariant geometric constraints (like relative position or visibility).

  • This allows for smart visual processing—using sight when it's clear, but ignoring misleading sight when it's distorted.

)4. Move Beyond Simple Prompting to Diverse, Condition-Specific Fine-Tuning

The limitations of structured prompting are that they introduce an answer/abstention trade-off that is model-dependent and not robust across different ambiguity types.

  • Implement a fine-tuning strategy using LoRA (Low-Rank Adaptation) on a diverse dataset containing both occlusion and perspective ambiguity scenarios (LoRA-Mixed approach).

  • This training regimen teaches the model to generalize the concept of unreliable observation rather than just memorizing how to follow a prompt.

  • The resulting system will possess a generalizable abstention capability, allowing it to reliably recognize when its input is fundamentally flawed, regardless of whether the flaw is occlusion or perspective distortion.

)What the Improved AI System Can Do (Specific Capabilities):

  1. Organized Robotic Navigation: An embodied agent can navigate a complex environment (e.g., a kitchen or warehouse) by not just looking at objects, but by assessing if its current view is reliable enough to determine object depth or size accurately. If perspective ambiguity arises from a tight corner, the system will proactively request a wider sweep of vision before attempting to grasp an item, preventing collisions and dropped objects.

  2. Advanced Inspection and Maintenance: A robotic inspection system can reliably identify subtle geometric differences (like the exact size comparison of two wall-mounted items) even when viewed from an oblique angle. It won't just give a potentially wrong answer; it will flag the specific viewpoint needed to confirm that measurement, ensuring high-precision quality control.

  3. Safety and Risk Assessment: In a dynamic environment, if sensor data is partially occluded (e.g., a safety barrier), the system can correctly abstain from making an immediate movement decision (like driving through or opening a door) until sufficient, reliable visual evidence is gathered from a known good viewpoint. This prevents dangerous overconfident actions based on incomplete data.

  4. Robust Human-Computer Interaction: In virtual reality interfaces, the AI can present information that is contextually honest. If a user asks Is this shelf taller than that one?, and the current camera angle makes the comparison ambiguous, the system will respond with a calibrated uncertainty statement (Cannot determine reliably from this view) rather than guessing incorrectly.

Sources

Related papers