InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
cs.CL, cs.CV, cs.HC
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: To be presented at EMNLP Main Conference
Code: https://github.com/maevehutch/insight
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of
Terminology
Abstract
Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. Existing benchmarks are predominantly constrained to static imagery and one-shot question answering and fail to capture the epistemic demands of this domain, where evidence is frequently occluded, distributed across linked views, or conditionally revealed through user agency. In this paper, we introduce InSight, a benchmark for agentic claim verification over interactive visualizations. The dataset consists of 21,349 claims derived from human-authored analytical narratives and grounded in fully interactive web-based environments. Agents must navigate these environments to determine whether a natural language claim is supported, refuted or not verifiable given the available evidence. Unlike traditional evaluations, InSight treats interaction traces as intrinsic proxies for reasoning, enabling a rigorous audit of how models seek and synthesize visual evidence. We evaluate state-of-the-art models, revealing that interactive verification remains a non-trivial challenge. We release InSight at https://github.com/maevehutch/insight.
Sources
- On Evaluation of Embodied Navigation Agents
- The BrowserGym Ecosystem for Web Agent Research
- Multimodal Web Navigation with Instruction-Finetuned Foundation Models
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
- Visual Test-time Scaling for GUI Agent Grounding
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- FEVER: a large-scale dataset for Fact Extraction and VERification
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering