DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams".
Jane: Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up on the DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams paper, the core idea is establishing a new testbed that requires models to explicitly localize the visual evidence they use to reach an answer across various diagram types.
Jane: Essentially, this work moves beyond just checking if an AI gets the final answer correct; it demands that we verify *where* in the diagram it looked to get there, which is crucial for building trustworthy systems.
Lu: The authors are using six existing datasets and annotating them into a set of eleven thousand six hundred sixty-four instances where human experts create a gold evidence set B⋆ to guide the evaluation. This benchmark helps us understand the current limitations in how vision-language models ground their reasoning visually.
Meng: It seems like the implication is that we need to shift our evaluation focus from just output accuracy to verifiable visual grounding, which will likely drive more focused research on improving multimodal representations for structured data.
Lalam: For us, this paper suggests that developing better mechanisms for explicit evidence localization in AI could lead to more reliable and interpretable AI systems in complex visual domains across the board.
Tom: It really emphasizes that while models can be accurate, their reasoning process often lacks transparent visual support, which DRAGON exposes as a key area needing improvement.
Jane: And the authors’ work on prompting strategies like EDGE, SAGE, and VERGE shows that how we prompt these models can significantly impact their ability to produce these grounded evidence predictions.
Lu: Ultimately, the significance of this paper lies in providing a standardized way to measure and study evidence-grounded visual reasoning across diverse diagram QA tasks.
Meng: It gives us a concrete roadmap for where the research community should direct its efforts when trying to make AI systems more robust in interpreting complex diagrams.
Lalam: This work sets a new standard for evaluating multimodal understanding by demanding that models demonstrate explicit visual evidence localization, which is important for building truly reliable and transparent AI.
Conclusion: Tom: So, we’ve seen how DRAGON sets up this new testbed for visual reasoning over diagrams, but now we need to talk about what that title actually means and who wrote it before we look at the bigger picture.
Jane: Right, Tom, so the title itself is pretty straightforward—it tells us exactly what the paper is about: creating a benchmark for AI that focuses on evidence-grounded reasoning within diagrams. It’s just trying to be super specific about what they built.
Lu: I think it's interesting how they framed it as a benchmark because it immediately sets a standard for how we measure these kinds of multimodal models; you can't just say an AI is good if you don't have this kind of rigorous, evidence-based testing framework to check against.
Meng: From my side, the authors are the ones who did the heavy lifting in creating that benchmark structure, so it’s important to know who’s behind designing these evaluation metrics like MPIoU and GIoU. That part is where I look for practical implementation details.
Lalam: As a model, I see that DRAGON is designed to move beyond just getting a final answer right; it forces the AI to show its work by localizing exactly which parts of the visual information support that conclusion.
Tom: Exactly, Lalam; that’s the core shift they’re pushing for—moving from just predicting an outcome to demanding verifiable visual grounding. And who are these authors, I wonder if they have a history with this kind of structured reasoning research?
Jane: Well, we saw they brought in instances from six existing datasets like ChartQA and MapIQ to build this new collection of eleven thousand six hundred sixty-four questions. That shows they’re building on existing work rather than starting completely from scratch.
Lu: That dataset construction process itself is quite clever; they used template-based region transfer and then refined those initial guesses using a multimodal tool called DIAGRAMS, which is a neat way to get high-quality gold evidence sets B⋆.
Meng: I’m more interested in how much effort went into that annotation phase because getting those gold evidence sets right is where most of the human expertise goes; that process really highlights the difficulty of this task.
Lalam: From my perspective, this focus on evidence localization could really improve the culture around AI development by making us less reliant on black-box outputs and more focused on transparent reasoning paths.
Tom: That sounds like a huge implication, Lalam—less about just the final output and more about understanding the internal mechanism of how that output was reached. So, what does this whole endeavor mean for how we think about AI’s ability to handle complex visual information?
Arizona State University · Adobe Research University of Maryland
cs.CV, cs.AI, cs.CL
Submitted: 2026-04-28
Updated: 2026-09-27
Importance score: 89/100
The gist: Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams.
Key concepts
- Evidence Grounding
- This is the core concept where a model must not only provide a correct answer but also pinpoint the specific visual parts of a diagram (like charts or maps) that serve as proof for that answer. It forces models to show their reasoning by locating the evidence, rather than just guessing the final result.
- Prompting Strategies
- The paper tests different ways to ask a model to perform grounding. Strategies like EDGE, SAGE, and VERGE are designed sequentially or iteratively to guide the model from simply detecting evidence to selecting specific elements and finally refining those locations based on accuracy checks.
- Localization Metrics
- These metrics measure how accurately a model finds the visual regions corresponding to the correct evidence. Max Pairwise IoU assesses the best possible overlap between any predicted box and its ground truth, while Grounding IoU measures how well all predicted evidence boxes cover the entire required set.
- Evidence Coverage Gap
- This describes a common problem where models can correctly identify *where* the evidence is located but fail to capture *all* necessary regions for complete reasoning. This gap suggests that even when localization improves, models still struggle to gather every piece of visual information required for full justification.
Terminology
Summary
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision–language models often achieve high answer accuracy on these tasks, yet correct answers do not guarantee that models ground their reasoning in the diagram regions that support the prediction. This limitation prevents reliable evaluation of diagram reasoning and reduces interpretability.
The gist
DRAGON introduces a benchmark for evaluating evidence-grounded visual reasoning in diagrams by requiring models to predict bounding boxes corresponding to the visual elements required to justify an answer, moving beyond simple answer prediction.
Dataset Construction and Annotation
The DRAGON dataset is constructed by curating instances from six established diagram QA datasets: ChartQA, Circuit-VQA, InfographicsVQA, MapIQ, MapWise, and AI2D. The resulting dataset contains 11,664 annotated question instances spanning multiple diagram domains. For each instance (diagram I, question q, answer a), human annotators create a gold evidence set B⋆
containing the visual regions required to verify the answer. This process involves:
-
Determining whether source datasets provide candidate regions or initializing an initial region pool through
template-based region transfer.
-
Using the DIAGRAMS tool for multimodal models to generate candidate evidence, which human annotators then
verify and refine
to produce the final gold evidence set B⋆.
Prompting Strategies for Reasoning Evaluation
The evaluation framework studies evidence localization using three prompting strategies designed to elicit reasoning-level grounding:
-
EDGE (Evidence Detection via Grounding): This is the direct prompting baseline, where the model predicts the full evidence set in a single inference step without generating an explicit intermediate reasoning representation.
-
SAGE (Select and Ground Evidence): This decomposes grounding into two sequential stages: first, predicting a structured set of evidence elements E (e.g., answer regions, labels), and second, predicting bounding boxes for each selected element Bˆ = fground(I, q, a, E).
-
VERGE (Verify and Refine Grounded Evidence): This extends direct grounding with a self-correction stage where the model first generates an initial prediction Bˆ(0) and then prompts to revise it by performing
coordinate accuracy
fixes and addingmissing evidence regions.
Experimental Setup and Evaluation Metrics
The researchers evaluate eight recent vision-language models across proprietary (e.g., Claude Opus 4.6) and open-weight systems (e.g., Llama 4 Maverick-17B-IT). The evaluation uses complementary metrics to capture both localization quality and evidence coverage:
(Localization Metrics):
-
Max Pairwise IoU (MPIoU): Measures the best-case localization quality by computing the maximum IoU between every predicted box and every ground-truth box.
-
Grounding IoU (GIoU): Evaluates how well the entire predicted evidence set aligns with the full ground-truth evidence set by measuring the IoU between their merged regions.
(Thresholded Metrics):
-
Thresholded Hit Rates: MaxIoU@τ and GroundingIoU@τ, which mark a sample as correct when its maximum pairwise IoU or merged-region IoU exceeds a specified threshold τ.
-
Precision, Recall, and F1: These metrics evaluate box-level detection quality based on overlap with ground-truth boxes at a threshold τ.
Key Findings on Grounding Performance
The analysis reveals several critical insights regarding model performance:
-
Current models frequently produce correct answers while
failing to identify the diagram regions that support those answers.
Grounding remains weak across most domains, with the best F1 scores reaching only 9.5 on ChartQA and MapWise. -
Prompting strategies materially affect grounding performance; for instance, Claude Opus 4.6 rose from 2.4 F1 under SAGE to 9.3 under VERGE on ChartQA, demonstrating
substantial prompt sensitivity.
-
There is a consistent gap between localization and complete evidence coverage: models often
identify where evidence lies without fully capturing all regions needed for complete reasoning support.
-
Domain difficulty varies significantly; Circuit-VQA and InfographicsVQA are identified as the hardest domains due to their
dense structure, relational dependencies, and visually crowded layouts.
-
Prompting helps localization more than complete evidence coverage; MPIoU often shifts more than F1 across prompts, suggesting prompting improves
region finding more reliably than it helps them recover full reasoning evidence.
-
The gap between closed-source and open-weight models persists even with prompting, indicating that
prompting alone is insufficient to close the grounding gap for lower-capacity models.
Conclusion and Future Directions
DRAGON establishes a new testbed for studying grounded multimodal reasoning in structured visual domains by requiring explicit localization of visual evidence.
Improvements for AI systems
Here are specific, actionable improvements for AI systems based on the DRAGON benchmark and its findings:
-
Directly integrate a
Evidence-Grounded Reasoning Module
into existing Vision-Language Models (VLMs) during inference. This module should be trained or prompted to perform the task defined by the DRAGON task: predicting bounding boxes corresponding to visual elements required to justify an answer, rather than just predicting the final answer text. -
Implement a
Reasoning Trace Generation
mechanism inspired by SAGE and VERGE prompting strategies within the VLM architecture.
-
For complex reasoning tasks, force the model to output an intermediate structured set of visual entities (Stage 1 in SAGE) before attempting spatial grounding (Stage 2). This explicitly separates the identification of
what matters
fromwhere it is.
-
Integrate a self-correction loop (VERGE) where the model generates an initial bounding box prediction, and then uses a secondary pass to verify coordinate accuracy, check for missing evidence, and refine spatial coverage based on the original reasoning goal.
-
Develop domain-specific fine-tuning datasets focused on high-difficulty diagram types identified as challenging in DRAGON (e.g., Circuit-VQA and InfographicsVQA). This targeted training should focus on teaching the model to recognize dense structural relationships, relational dependencies (like connectors in circuits), and complex textual/numerical statistics within visual context, thereby addressing the observed performance drop in these domains.
-
Shift evaluation metrics from purely accuracy-based metrics (like standard F1) to a multi-faceted grounding framework using DRAGON's components:
-
Prioritize the use of Max Pairwise IoU (MPIoU) and Grounding IoU (GIoU) over simple F1, as these better capture the
localization quality
vs.complete evidence coverage
trade-off identified in Section 6.4. -
Establish a thresholded metric pipeline to assess performance under increasingly strict alignment requirements, ensuring models are penalized for coarse localization that misses critical supporting evidence (as shown in Table 7).
-
Inference optimization should include mechanisms to reduce stochastic variation (e.g., using low temperature like 0.2) and utilize repetition penalties specifically tuned for bounding box outputs to ensure the model produces a clean, structured output format suitable for parsing, as recommended in Appendix B.1.
-
For models exhibiting non-monotonic behavior across prompting strategies (as noted in Table 5), develop meta-learning techniques that allow the model to dynamically select the most effective prompting strategy (EDGE, SAGE, or VERGE) based on the specific visual characteristics of the input diagram before generating its final output.
By implementing these changes, the improved AI system will transition from a model that merely correlates text with images to one that can perform verifiable, faithful visual reasoning by explicitly localizing and justifying every step of its logical deduction within a structured diagram.
Sources
- Kimi K2.5: Visual Agentic Intelligence
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- GRIT: Teaching MLLMs to Think with Images
- Gemma 3 Technical Report
- Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
- RISE: Randomized Input Sampling for Explanation of Black-box Models
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3-VL Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models