EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
cs.CL, cs.AI
Submitted: 2026-01-06
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 Main Conference. 35 pages, 5 figures, 31 tables
Code: https://github.com/myweiii/EpiQAL
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level.
Terminology
Abstract
Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, to our knowledge the first diagnostic benchmark for epidemiological question answering over research literature, comprising three subsets built from open-access articles across diverse diseases. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.
Sources
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- Adversarial Filters of Dataset Biases
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- The Llama 3 Herd of Models
- Measuring Massive Multitask Language Understanding
- Mistral 7B
- Dynabench: Rethinking Benchmarking in NLP
- Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
- Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- GPT-4 Technical Report
- emrQA: A Large Corpus for Question Answering on Electronic Medical Records
- OpenAI GPT-5 System Card
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals
- Utilizing Large Language Models for Zero-Shot Medical Ontology Extension from Clinical Notes
- WebDancer: Towards Autonomous Information Seeking Agency
- KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMs
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering