Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
cs.CV, cs.AI
Submitted: 2025-06-11
Updated: 2026-08-30
Comments: Accepted by IJCV
DOI: 10.1007/s11263-026-02980-3
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse
Terminology
Abstract
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- A-BDD: Leveraging Data Augmentations for Safe Autonomous Driving in Adverse Weather and Lighting
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
- Scaling Out-of-Distribution Detection for Real-World Settings
- DAWN: Vehicle Detection in Adverse Weather Nature Dataset
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
- Dolphins: Multimodal Language Model for Driving
- Open-World Object Manipulation using Pre-trained Vision-Language Models
- NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models