DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
cs.CL
Submitted: 2025-08-20
Updated: 2026-09-28
Journal ref: EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information seeking: synthesizing multimodal evidence
Terminology
Abstract
Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information seeking: synthesizing multimodal evidence scattered across multiple documents and structural formats. Existing QAs remain narrow in scope, relying on unimodal text and short-span reasoning that fail to capture the complexity of real information-seeking. We introduce DocHop-QA, a benchmark of 11,379 instances for evaluating multimodal, multi-document, multi-hop scientific QA. Built from publicly available PubMed articles, DocHop-QA incorporates textual passages, tables, and layout cues, enabling cross-document inference without explicit hyperlinks. To scale realistic QA construction, we develop an LLM-driven generation pipeline grounded in 11 scientific reasoning concepts, producing diverse and coherent question-answer pairs. To highlight the utility and versatility of the dataset, we propose a task-driven evaluation framework spanning four settings, including generative answering, multimodal evidence integration and structured index prediction. Experiments show that current models struggle with DocHop-QA's long-context, multi-evidence demands, establishing it as a rigorous testbed for advancing next-generation scientific QA systems.
Sources
- Longformer: The Long-Document Transformer
- Multi-hop Question Answering via Reasoning Chains
- HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data
- FORTAP: Using Formulas for Numerical-Reasoning-Aware Table Pretraining
- PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
- LoRA: Low-Rank Adaptation of Large Language Models
- Qwen2.5-VL Technical Report
- BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain
- SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
- Exploiting Reasoning Chains for Multi-hop Science Question Answering
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- BERTScore: Evaluating Text Generation with BERT
- FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
- Mitigating Lost-in-Retrieval Problems in Retrieval Augmented Multi-Hop Question Answering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering