When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
cs.AI, cs.CV
Submitted: 2026-08-25
Updated: 2026-08-30
Comments: EMNLP 2026 Main
Code: https://github.com/jaaack-wang/benchmarking-interative-visual-grounding
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target.
Terminology
Abstract
Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions are often incomplete or ambiguous, requiring participants to establish shared understanding through interaction. We introduce a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), varying how much target information is provided upfront and how much must be acquired through dialogue. Across four human-grounded visual contexts and four interaction protocols, current LVLMs perform significantly below task-level human baselines. Interaction can help when follow-up questions refine or repair an initial target description. Performance is lowest when no initial description is provided and target information must be acquired through questions, indicating that proactive question-driven grounding remains difficult. LVLMs are also poorly calibrated, often reporting confidence that exceeds their empirical accuracy. Follow-up studies confirm these patterns across varied description sources (human versus AI), reasoning efforts, repeated interactions, description providers, and visual contexts. Overall, interactive visual grounding remains challenging, requiring visual matching, information seeking and synthesis.
Sources
- Qwen3-VL Technical Report
- Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
- Gemini: A Family of Highly Capable Multimodal Models
- LLMs and people both learn to form conventions -- just not with each other
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- GPT-4 Technical Report
- Modeling Context in Referring Expressions
- LVLMs and Humans Ground Differently in Referential Communication
- InViG: Benchmarking Interactive Visual Grounding with 500K Human-Robot Interactions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection