Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

arXiv:2608.10954 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

University of Chinese Academy of Sciences · Tencent CDG · Institute of Information Engineering, Chinese Academy of Sciences

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-27

Comments: Accepted by IJCV

Code: https://github.com/hiyouga/EasyR1

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: The paper introduces AD2-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in visually adverse and complex urban scenes, and EGVOR (Evidence-Grounded Visual

Terminology

Summary

The paper introduces AD2-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in visually adverse and complex urban scenes, and EGVOR (Evidence-Grounded Visual Reasoning), a method to improve trustworthy reasoning in such conditions.

Key contributions and findings:

  1. AD2-Bench Benchmark: A comprehensive benchmark with approximately 10K images and 70K QA pairs, covering diverse environmental degradations (rain, fog, snow, night, etc.) and complex layouts. It features a Hierarchical Visual Diagnosis that decomposes reasoning into a structured Chain of Evidence (CoE), enabling a fine-grained, trustworthiness-oriented evaluation framework with metrics like Hallucination Resistance Score (HRS), Logical Reliability Score (LRS), Self-Consistency Score (SCS), and Explainability Alignment Score (EAS).

  2. Probabilistic Bottleneck Analysis: The paper theoretically diagnoses reasoning failures in adverse scenarios as two primary bottlenecks: Spatial Ambiguity (failing to distinguish targets from background clutter) and Semantic Uncertainty (misinterpreting semantics due to feature degradation). This analysis shows that robust cognition is predicated on precise evidence acquisition.

  3. EGVOR Methodology: To address these bottlenecks, EGVOR reformulates multimodal reasoning from implicit text generation into the explicit, sequential construction of Evidence Atoms—structured triplets ⟨bt, ht, ct⟩ (Spatial Anchor, Region-Aware Latent State, Grounded Description) that enforce strict spatial-semantic alignment. It uses a two-stage hierarchical curriculum:

  • Stage I (Reflective SFT): Establishes the syntactic structure of evidence chains and includes error-induced self-correction.

  • Stage II (Cognitive Alignment via RL): Uses Group Relative Policy Optimization (GRPO) with a composite reward function (Jtotal) that includes a Hierarchical Robust Gaussian-Wasserstein (HRGW) spatial reward (Ψspatial), a semantic alignment reward (Ξalign), and a cognitive path diversity reward (Λpath).

  1. Experimental Results:
  • On AD2-Bench, EGVOR (7B) achieves a state-of-the-art average score of 67.66%, significantly outperforming its 10x larger counterpart Qwen2.5-VL-72B (60.67%) and other baselines.

  • EGVOR demonstrates strong generalization on external benchmarks like V* Bench, HR-Bench, MME-RealWorld, and TreeBench, showing improvements in fine-grained perception, robustness, and grounded reasoning.

  • Ablation studies confirm the importance of each component, particularly the spatial constraint (Ψspatial), which is pivotal for improving localization (mIoU) and overall reasoning accuracy.

  • The model exhibits improved robustness against visual disturbances, with a lower Semantic Entropy Gap (∆) between clear and adverse conditions (0.278 vs. 0.505 for standard CoT).

  • EGVOR is also more efficient, achieving higher performance with fewer tokens (144 vs. 212) compared to its SFT stage, resulting in a much higher Marginal Utility (51.7 vs. 5.0).

The paper concludes that robust multimodal cognition in complex real-world scenarios is predicated on the verifiable acquisition of visual evidence, and EGVOR provides a robust framework for trustworthy and interpretable decision-making.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

1. Implement Evidence-Atom-Based Reasoning for Vision-Language Models

  • Replace implicit text generation with structured triplets ⟨spatial anchor, region-aware latent state, grounded description⟩ for all visual QA tasks.

  • This forces the model to explicitly localize and describe objects before answering, reducing hallucination in cluttered scenes.

2. Add a Hierarchical Visual Diagnosis Module

  • Decompose any visual reasoning task into a Chain of Evidence with stages: detect → localize → describe → infer → conclude.

  • Use the paper’s metrics (HRS, LRS, SCS, EAS) as internal reward signals during training, not just evaluation, to penalize ungrounded answers.

3. Apply Two-Stage Curriculum Learning for Robustness

  • Stage 1: Train on error-induced self-correction examples (where the model sees a wrong answer, then the correct evidence chain) to learn syntactic structure of reasoning.

  • Stage 2: Use GRPO with a composite reward that includes:

  • Spatial reward (Ψspatial) via Hierarchical Robust Gaussian-Wasserstein distance to penalize mislocalization.

  • Semantic alignment reward (Ξalign) to match evidence descriptions to ground truth.

  • Path diversity reward (Λpath) to encourage multiple valid reasoning paths, improving self-consistency.

4. Inject Probabilistic Bottleneck-Aware Training Data

  • During fine-tuning, deliberately add images with spatial ambiguity (targets near background clutter) and semantic uncertainty (degraded features from rain/fog/night).

  • Train the model to explicitly output uncertainty flags when evidence is insufficient, rather than forcing a confident guess.

5. Use Token-Efficient Evidence Chains

  • Optimize the model to produce compact evidence atoms (144 tokens vs. 212) by training on minimal-but-sufficient descriptions, improving inference speed and marginal utility.

The improved AI system can:

  • Answer visual questions in adverse conditions (fog, night, rain) with 67.66% accuracy at 7B scale, outperforming 72B models.

  • Provide interpretable reasoning: each answer is backed by a verifiable chain of spatial anchors and grounded descriptions.

  • Self-correct errors by re-examining evidence atoms when initial reasoning fails.

  • Maintain consistent performance across clear and degraded inputs (semantic entropy gap reduced by 45%).

  • Handle fine-grained localization tasks (e.g., which object is behind the car?) with higher mIoU and fewer hallucinations.

  • Generalize to unseen benchmarks (V* Bench, HR-Bench, MME-RealWorld) without retraining, showing robust transfer to real-world perception tasks.

Sources

Related papers