DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
summary
The gist
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams.
In short
DRAGON introduces a benchmark for evaluating evidence-grounded visual reasoning in diagrams by requiring models to predict bounding boxes for supporting visual elements, moving beyond just answering questions. The dataset combines data from six sources and uses three prompting strategies to test how well models localize the exact regions needed to justify an answer.
Key concepts
- Evidence Grounding
- This is the core concept where a model must not only provide a correct answer but also pinpoint the specific visual parts of a diagram (like charts or maps) that serve as proof for that answer. It forces models to show their reasoning by locating the evidence, rather than just guessing the final result.
- Prompting Strategies
- The paper tests different ways to ask a model to perform grounding. Strategies like EDGE, SAGE, and VERGE are designed sequentially or iteratively to guide the model from simply detecting evidence to selecting specific elements and finally refining those locations based on accuracy checks.
- Localization Metrics
- These metrics measure how accurately a model finds the visual regions corresponding to the correct evidence. Max Pairwise IoU assesses the best possible overlap between any predicted box and its ground truth, while Grounding IoU measures how well all predicted evidence boxes cover the entire required set.
- Evidence Coverage Gap
- This describes a common problem where models can correctly identify *where* the evidence is located but fail to capture *all* necessary regions for complete reasoning. This gap suggests that even when localization improves, models still struggle to gather every piece of visual information required for full justification.
Terminology used across episodes
This episode discusses
- DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams · Paper Radio
- Kimi K2.5: Visual Agentic Intelligence
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- GRIT: Teaching MLLMs to Think with Images
- Gemma 3 Technical Report
- Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
- RISE: Randomized Input Sampling for Explanation of Black-box Models
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3-VL Technical Report
The paper
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams · Read on arXiv
Arizona State University · Adobe Research University of Maryland
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams".
Jane: Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up on the DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams paper, the core idea is establishing a new testbed that requires models to explicitly localize the visual evidence they use to reach an answer across various diagram types.
Jane: Essentially, this work moves beyond just checking if an AI gets the final answer correct; it demands that we verify *where* in the diagram it looked to get there, which is crucial for building trustworthy systems.
Lu: The authors are using six existing datasets and annotating them into a set of eleven thousand six hundred sixty-four instances where human experts create a gold evidence set B⋆ to guide the evaluation. This benchmark helps us understand the current limitations in how vision-language models ground their reasoning visually.
Meng: It seems like the implication is that we need to shift our evaluation focus from just output accuracy to verifiable visual grounding, which will likely drive more focused research on improving multimodal representations for structured data.
Lalam: For us, this paper suggests that developing better mechanisms for explicit evidence localization in AI could lead to more reliable and interpretable AI systems in complex visual domains across the board.
Tom: It really emphasizes that while models can be accurate, their reasoning process often lacks transparent visual support, which DRAGON exposes as a key area needing improvement.
Jane: And the authors’ work on prompting strategies like EDGE, SAGE, and VERGE shows that how we prompt these models can significantly impact their ability to produce these grounded evidence predictions.
Lu: Ultimately, the significance of this paper lies in providing a standardized way to measure and study evidence-grounded visual reasoning across diverse diagram QA tasks.
Meng: It gives us a concrete roadmap for where the research community should direct its efforts when trying to make AI systems more robust in interpreting complex diagrams.
Lalam: This work sets a new standard for evaluating multimodal understanding by demanding that models demonstrate explicit visual evidence localization, which is important for building truly reliable and transparent AI.
Conclusion: Tom: So, we’ve seen how DRAGON sets up this new testbed for visual reasoning over diagrams, but now we need to talk about what that title actually means and who wrote it before we look at the bigger picture.
Jane: Right, Tom, so the title itself is pretty straightforward—it tells us exactly what the paper is about: creating a benchmark for AI that focuses on evidence-grounded reasoning within diagrams. It’s just trying to be super specific about what they built.
Lu: I think it's interesting how they framed it as a benchmark because it immediately sets a standard for how we measure these kinds of multimodal models; you can't just say an AI is good if you don't have this kind of rigorous, evidence-based testing framework to check against.
Meng: From my side, the authors are the ones who did the heavy lifting in creating that benchmark structure, so it’s important to know who’s behind designing these evaluation metrics like MPIoU and GIoU. That part is where I look for practical implementation details.
Lalam: As a model, I see that DRAGON is designed to move beyond just getting a final answer right; it forces the AI to show its work by localizing exactly which parts of the visual information support that conclusion.
Tom: Exactly, Lalam; that’s the core shift they’re pushing for—moving from just predicting an outcome to demanding verifiable visual grounding. And who are these authors, I wonder if they have a history with this kind of structured reasoning research?
Jane: Well, we saw they brought in instances from six existing datasets like ChartQA and MapIQ to build this new collection of eleven thousand six hundred sixty-four questions. That shows they’re building on existing work rather than starting completely from scratch.
Lu: That dataset construction process itself is quite clever; they used template-based region transfer and then refined those initial guesses using a multimodal tool called DIAGRAMS, which is a neat way to get high-quality gold evidence sets B⋆.
Meng: I’m more interested in how much effort went into that annotation phase because getting those gold evidence sets right is where most of the human expertise goes; that process really highlights the difficulty of this task.
Lalam: From my perspective, this focus on evidence localization could really improve the culture around AI development by making us less reliant on black-box outputs and more focused on transparent reasoning paths.
Tom: That sounds like a huge implication, Lalam—less about just the final output and more about understanding the internal mechanism of how that output was reached. So, what does this whole endeavor mean for how we think about AI’s ability to handle complex visual information?
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization