BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding
cs.AI
Submitted: 2026-01-08
Updated: 2026-08-28
Comments: Accepted to Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cross-disciplinary communication.
Terminology
Abstract
Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cross-disciplinary communication. Two challenges, High Information Density (HID) and Multi-Step Reasoning (MSR), pose unique difficulties for precise automatic experimental understanding. Extracting structured knowledge, e.g., Knowledge Graphs (KGs), is an effective approach to address the HID and MSR. However, existing biomedical datasets for structured knowledge Information Extraction (IE) are limited to a general or coarse-grained level, hindering fine-grained experimental understanding. To address this gap, we introduce Biomedical Protocol Information Extraction Dataset (BioPIE), a dataset providing procedure-centric KGs that captures entities, actions, and relations at a scale sufficient for reasoning across biomedical protocols. We evaluate both supervised and LLM-based IE methods on BioPIE to verify its effectiveness, and implement a biomedical question answering system to provide a quantitative illustration of BioPIE's effectiveness for downstream understanding tasks. The experimental results demonstrate improved understanding performance on both the HID and MSR question sets.
Sources
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
- ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection