Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
cs.AI
Submitted: 2026-09-10
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses.
Terminology
Abstract
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents
Sources
- SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
- M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
- Towards Long-horizon Agentic Multimodal Search
- Accelerating scientific discovery with Co-Scientist
- MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
- WebSailor: Navigating Super-human Reasoning for Web Agent
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
- MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection