Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
University of Chinese Academy of Sciences · Tencent CDG · Institute of Information Engineering, Chinese Academy of Sciences
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-27
Comments: Accepted by IJCV
Code: https://github.com/hiyouga/EasyR1
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper introduces AD2-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in visually adverse and complex urban scenes, and EGVOR (Evidence-Grounded Visual
Terminology
Summary
The paper introduces AD2-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in visually adverse and complex urban scenes, and EGVOR (Evidence-Grounded Visual Reasoning), a method to improve trustworthy reasoning in such conditions.
Key contributions and findings:
-
AD2-Bench Benchmark: A comprehensive benchmark with approximately 10K images and 70K QA pairs, covering diverse environmental degradations (rain, fog, snow, night, etc.) and complex layouts. It features a
Hierarchical Visual Diagnosis
that decomposes reasoning into a structuredChain of Evidence
(CoE), enabling a fine-grained, trustworthiness-oriented evaluation framework with metrics like Hallucination Resistance Score (HRS), Logical Reliability Score (LRS), Self-Consistency Score (SCS), and Explainability Alignment Score (EAS). -
Probabilistic Bottleneck Analysis: The paper theoretically diagnoses reasoning failures in adverse scenarios as two primary bottlenecks: Spatial Ambiguity (failing to distinguish targets from background clutter) and Semantic Uncertainty (misinterpreting semantics due to feature degradation). This analysis shows that robust cognition is predicated on precise evidence acquisition.
-
EGVOR Methodology: To address these bottlenecks, EGVOR reformulates multimodal reasoning from implicit text generation into the explicit, sequential construction of
Evidence Atoms
—structured triplets ⟨bt, ht, ct⟩ (Spatial Anchor, Region-Aware Latent State, Grounded Description) that enforce strict spatial-semantic alignment. It uses a two-stage hierarchical curriculum:
-
Stage I (Reflective SFT): Establishes the syntactic structure of evidence chains and includes error-induced self-correction.
-
Stage II (Cognitive Alignment via RL): Uses Group Relative Policy Optimization (GRPO) with a composite reward function (Jtotal) that includes a Hierarchical Robust Gaussian-Wasserstein (HRGW) spatial reward (Ψspatial), a semantic alignment reward (Ξalign), and a cognitive path diversity reward (Λpath).
- Experimental Results:
-
On AD2-Bench, EGVOR (7B) achieves a state-of-the-art average score of 67.66%, significantly outperforming its 10x larger counterpart Qwen2.5-VL-72B (60.67%) and other baselines.
-
EGVOR demonstrates strong generalization on external benchmarks like V* Bench, HR-Bench, MME-RealWorld, and TreeBench, showing improvements in fine-grained perception, robustness, and grounded reasoning.
-
Ablation studies confirm the importance of each component, particularly the spatial constraint (Ψspatial), which is pivotal for improving localization (mIoU) and overall reasoning accuracy.
-
The model exhibits improved robustness against visual disturbances, with a lower Semantic Entropy Gap (∆) between clear and adverse conditions (0.278 vs. 0.505 for standard CoT).
-
EGVOR is also more efficient, achieving higher performance with fewer tokens (144 vs. 212) compared to its SFT stage, resulting in a much higher Marginal Utility (51.7 vs. 5.0).
The paper concludes that robust multimodal cognition in complex real-world scenarios is predicated on the verifiable acquisition of visual evidence, and EGVOR provides a robust framework for trustworthy and interpretable decision-making.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
1. Implement Evidence-Atom-Based Reasoning for Vision-Language Models
-
Replace implicit text generation with structured triplets ⟨spatial anchor, region-aware latent state, grounded description⟩ for all visual QA tasks.
-
This forces the model to explicitly localize and describe objects before answering, reducing hallucination in cluttered scenes.
2. Add a Hierarchical Visual Diagnosis Module
-
Decompose any visual reasoning task into a
Chain of Evidence
with stages: detect → localize → describe → infer → conclude. -
Use the paper’s metrics (HRS, LRS, SCS, EAS) as internal reward signals during training, not just evaluation, to penalize ungrounded answers.
3. Apply Two-Stage Curriculum Learning for Robustness
-
Stage 1: Train on error-induced self-correction examples (where the model sees a wrong answer, then the correct evidence chain) to learn syntactic structure of reasoning.
-
Stage 2: Use GRPO with a composite reward that includes:
-
Spatial reward (Ψspatial) via Hierarchical Robust Gaussian-Wasserstein distance to penalize mislocalization.
-
Semantic alignment reward (Ξalign) to match evidence descriptions to ground truth.
-
Path diversity reward (Λpath) to encourage multiple valid reasoning paths, improving self-consistency.
4. Inject Probabilistic Bottleneck-Aware Training Data
-
During fine-tuning, deliberately add images with spatial ambiguity (targets near background clutter) and semantic uncertainty (degraded features from rain/fog/night).
-
Train the model to explicitly output
uncertainty flags
when evidence is insufficient, rather than forcing a confident guess.
5. Use Token-Efficient Evidence Chains
- Optimize the model to produce compact evidence atoms (144 tokens vs. 212) by training on minimal-but-sufficient descriptions, improving inference speed and marginal utility.
The improved AI system can:
-
Answer visual questions in adverse conditions (fog, night, rain) with 67.66% accuracy at 7B scale, outperforming 72B models.
-
Provide interpretable reasoning: each answer is backed by a verifiable chain of spatial anchors and grounded descriptions.
-
Self-correct errors by re-examining evidence atoms when initial reasoning fails.
-
Maintain consistent performance across clear and degraded inputs (semantic entropy gap reduced by 45%).
-
Handle fine-grained localization tasks (e.g.,
which object is behind the car?
) with higher mIoU and fewer hallucinations. -
Generalize to unseen benchmarks (V* Bench, HR-Bench, MME-RealWorld) without retraining, showing robust transfer to real-world perception tasks.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models