BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions".
Jane: BioMol-MQA is a novel question-answering dataset designed to test and improve Large Language Model (LLM) reasoning capabilities over complex, multi-modal bio-molecular interactions.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We've talked a lot about how BioMol-MQA is structured as a dataset and why it’s designed to test the limits of current RAG systems by demanding reasoning across different data types. So, let's touch on the title and authors again as we wrap up this discussion.
Jane: It’s worth remembering that the paper, "BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions," was put out by Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, and Suhang Wang from Pennsylvania State University.
Lu: That team has built a really solid foundation here for testing how complex AI can handle the intersection of chemistry and clinical knowledge. Their work on creating this structured dataset is significant for setting new standards in how we evaluate these systems.
Meng: The implication we discussed earlier comes back to practical deployment; if the models can successfully navigate these multi-modal interactions, it means we could move closer to AI tools that can give nuanced advice regarding drug combinations safely.
Lalam: For the wider field, this signals a clear direction for developing more sophisticated knowledge synthesis capabilities in AI, moving past single-source limitations toward truly comprehensive understanding of complex domains like polypharmacy.
Tom: So, to summarize this whole discussion about BioMol-MQA, it’s about creating a rigorous environment where LLMs have to demonstrate that they can pull information from both molecular structures and textual descriptions to solve difficult questions.
Jane: It really highlights that the future of effective RAG systems isn't just about better retrieval, but about building architectures that can effectively reason over these diverse modalities when the context demands it.
Lu: And this dataset provides a concrete benchmark for measuring that exact kind of integrated reasoning ability in bio-molecular contexts.
Meng: From an engineering view, it confirms that the complexity is real and requires a more robust pipeline than what we currently deploy for these types of queries.
Lalam: Ultimately, the value here is in pushing the entire ecosystem toward building AI that can synthesize knowledge from different domains simultaneously to tackle high-stakes problems effectively.
Conclusion: Tom: So we've been diving deep into how BioMol-MQA sets up these LLMs to tackle complex bio-molecular reasoning, and now it's time for us to wrap up this segment by talking about what this whole dataset is actually titled and who put it together.
Jane: That’s right, we’ve covered the technical details of the dataset construction, but let’s focus on the core idea behind "BioMol-MQA" itself and why those authors chose to put it out into the public domain.
Lu: The title tells you everything; it signals that they're moving beyond just looking at text or just looking at structures, aiming for a unified understanding of how drugs and proteins actually interact.
Meng: From an engineering standpoint, having a standardized dataset like this is huge because it gives us something concrete to test our retrieval frameworks against when we build the next generation of reasoning systems.
Lalam: I see the core implication is that we’re forcing AI models to develop a more holistic way of thinking about biological systems, which could eventually mean better decision support in healthcare settings.
Tom: Exactly, and when you look at the authors, you see a team clearly dedicated to bridging those different scientific worlds—chemistry and clinical knowledge—which is really impressive.
Jane: I think what’s important to grasp for our listeners is that this isn't just another data dump; it’s a carefully constructed challenge designed specifically to find the limits of how much reasoning an AI can do across different data types at the same time.
Lu: That multi-modal structure means they are testing if an AI can seamlessly switch between understanding a chemical formula, reading a clinical summary, and knowing what those two things mean when put together.
Meng: And for practical impact, this means that as we build these systems for real-world use, we need to ensure our RAG pipelines can handle this kind of integrated complexity without breaking down.
Lalam: I feel like the most significant vision here is how this kind of data forces us to build AI culture around synthesis rather than just simple pattern matching, which is a big step forward for how we think about autonomous systems.
Tom: So, as we look at the whole picture, BioMol-MQA isn't just a dataset; it’s a blueprint for the next level of sophisticated AI that needs to truly grasp complex biological realities. What do you all think this means moving forward?
Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang
Pennsylvania State University
cs.CL
Submitted: 2025-06-06
Updated: 2026-10-01
Code: https://github.com/LSYS/lexicalrichness
Importance score: 84/100
The gist: BioMol-MQA is a novel question-answering dataset designed to test and improve Large Language Model (LLM) reasoning capabilities over complex, multi-modal bio-molecular interactions.
Key concepts
- Multimodal Knowledge Graph (KG)
- This is a structured database linking drugs and proteins as nodes, with edges showing their interactions. It includes text about the entities and molecular structure data (SMILES) to provide rich, multi-faceted information for complex queries.
- Polypharmacy
- This refers to the use of multiple medications by a patient. The dataset focuses on this area because understanding how different drugs interact (Drug-Drug Interactions or DDI) is critical in healthcare and requires deep reasoning.
- Multi-hop Reasoning
- This is when an LLM must connect information across multiple steps or paths within the knowledge graph to find an answer. For example, it might need to follow a chain of interactions between three different entities to solve a single question.
- Retrieval Augmented Generation (RAG)
- RAG is a technique where an LLM first retrieves relevant, specific information from an external knowledge base before generating an answer. BioMol-MQA tests if LLMs can effectively use this retrieval step to handle complex scientific questions.
Terminology
Summary
BioMol-MQA is a novel question-answering dataset designed to test and improve Large Language Model (LLM) reasoning capabilities over complex, multi-modal bio-molecular interactions. This dataset addresses the gap in existing RAG systems that focus on single modalities by integrating a multimodal knowledge graph (KG) with text and molecular structures to create challenging questions for LLMs. The work is crucial because it signals the necessity for strong Retrieval Augmented Generation (RAG) frameworks capable of retrieving and synthesizing diverse domain-specific knowledge, such as polypharmacy information, which is vital in high-stakes fields like healthcare.
The Gist
BioMol-MQA is a new question-answering (QA) dataset on polypharmacy, which is composed of two parts (i) a multimodal knowledge graph (KG) with text and molecular structure for information retrieval; and (ii) challenging questions that designed to test LLM capabilities in retrieving and reasoning over multimodal KG to answer questions.
Dataset Composition and Modalities
The BioMol-MQA dataset is structured around three core modalities:
-
A knowledge graph of drugs and proteins, where nodes are drugs or proteins, and edges represent interactions between entities (Drug-Drug Interactions (DDI) and Drug-Protein Interactions (DPI)). The graph contains 494 unique drugs, 198 proteins, 18.5K drug-drug edges, and 314 drug-protein edges.
-
Free text associated with each node in the graph, providing background knowledge such as uses or behaviors of the entities. This text is sourced from Wikipedia summaries and PubMed abstracts.
-
Molecular structure data of drugs represented by SMILES (Simplified Molecular Input Line Entry System) strings, which provide insight into molecular composition and potential interactions between atoms.
Data Construction Pipeline
The dataset development follows a five-stage pipeline:
-
Base Data Collection: Assembling the knowledge graph and augmenting its nodes with text and molecular structures. Entity names are resolved by querying various databases (PubChem, ChemSpider, STRING) to obtain generic or commercial names for drugs and protein names from gene IDs. DPI labels are derived by querying the STITCH database for interaction modes (e.g., binding/inhibition).
-
Text Post-Processing: Source texts are post-processed using an LLM (GPT-4o) to rewrite them, ignoring historical data and using domain-specific jargon to enhance question complexity, aiming for a
semantically dense corpus.
-
Molecular Interaction Extraction: GPT-4o is prompted with SMILES strings of two drugs to describe potential molecular associations, yielding labels such as
Hydrogen Bonding,
Steric Clashes,
orIonic Interaction.
This transforms the base graph into a multi-graph. -
Question Generation: LLMs (GPT-4.1) are prompted to create questions by integrating relational information with background data (text or SMILES). Questions are designed to test knowledge retrieval from at least two modality combinations, such as
G + T or G + S.
The task is formally defined as a function F: (D, Q) → N. -
Question Verification: Quality is assessed using both automatic verification (LLM-as-a-judge with Claude-3.7 Sonnet) and human evaluation, utilizing a rubric with metrics including Clarity, Coverage, Assumptions, and Inferable answer.
Question Design and Complexity
The questions are constructed to require multi-modal reasoning. The task definition requires the LLM to utilize at least two modality combinations (G + T or G + S) to solve a query Q.
Questions are not multiple-choice; they are open-ended, requiring retrieval and synthesis. The dataset includes both single-hop questions (involving two connected entities) and multi-hop reasoning questions involving paths longer than one edge. Quantitative metrics analyzed include Question Length (average 66.6 tokens), Type-to-Token Ratio (TTR of 0.84), Shannon Entropy (average 5.67 bits), and Dependency Parse Tree Depth (average 10.86).
Experimental Findings and Limitations
Experimental results on BioMol-MQA show that current LLMs are inept at solving these questions
without background data, but their performance spikes
when provided with the necessary context, confirming the necessity of RAG. Benchmarking against seven models revealed that while zero-shot performance is low (average EM/F1/BERTScore 0.22/0.28/0.77), upper-bound performance with gold data is significantly better (average 0.62/0.67/0.83). Retrieval benchmarks indicate that simple BM25 outperforms dense embedding models for text retrieval, and basic database retrievers like Neo4j outperform trained GNN retrievers due to the graph's size constraints.
Improvements for AI systems
Here are specific improvements for AI systems based on the BioMol-MQA dataset and methodology:
-
Enhanced Multi-Modal Retrieval Augmented Generation (RAG) Frameworks:
-
Improved Domain-Specific Knowledge Grounding in RAG Systems:
-
Development of Specialized Modality Adapters for LLMs:
-
Creation of a Synthetic Data Generation Pipeline for Complex Reasoning Tasks:
- Enhanced Multi-Modal Retrieval Augmented Generation (RAG) Frameworks:
The system can be improved by moving beyond single-modality retrieval to a robust, integrated multi-modal RAG pipeline. This involves developing hybrid retrievers that simultaneously query knowledge graphs (for relational data), text corpora (for background context), and molecular structure databases (via SMILES encoding).
The improved AI system will be able to:
-
Retrieve answers by synthesizing information across different formats—e.g., identifying a drug's interaction with a protein (KG + Text) while simultaneously checking the chemical mechanism via its structure (SMILES).
-
Perform complex, multi-hop reasoning that requires connecting disparate pieces of information from graph structures, textual background, and molecular properties to arrive at a final answer.
- Improved Domain-Specific Knowledge Grounding in RAG Systems:
The system can be improved by specifically grounding its responses in high-quality, domain-specific knowledge derived from the BioMol-MQA KG (drug-drug/protein interactions). This moves LLMs from general knowledge to expert reasoning within a clinical/pharmacological context.
The improved AI system will be able to:
-
Provide medically accurate answers regarding polypharmacy and drug interactions, minimizing hallucinations by requiring the LLM to explicitly cite or reason over retrieved graph triples and chemical evidence (SMILES features).
-
Handle queries that require understanding the
why
behind an interaction (e.g., identifying specific molecular mechanisms like Hydrogen Bonding or Steric Clashes) rather than just naming a drug.
- Development of Specialized Modality Adapters for LLMs:
The system can be improved by training or fine-tuning LLMs with modality-specific adapters that allow them to efficiently utilize the different input types (Graph, Text, SMILES) during the reasoning process. This addresses the limitation where current models struggle to jointly utilize these sources effectively.
The improved AI system will be able to:
-
Effectively map abstract concepts in a question onto the appropriate data structure (e.g., recognizing when a query requires graph traversal vs. molecular property lookup).
-
Perform specialized tasks, such as molecular interaction extraction (using SMILES), with high fidelity, by leveraging LLMs trained on chemical reasoning tasks or fine-tuned for structure analysis.
- Creation of a Synthetic Data Generation Pipeline for Complex Reasoning Tasks:
The system can be improved by implementing the synthetic data generation pipeline described in Section 3.2 and 3.4 to create highly complex, challenging questions that test the limits of multimodal reasoning, rather than relying on existing template-based datasets (like MoleculeQA).
The improved AI system will be able to:
-
Be rigorously tested against frontier LLMs (like GPT-4o or Claude 3.7 Sonnet) to determine its true capability for complex, multi-modal synthesis and inference in high-stakes scenarios.
-
Be fine-tuned on a diverse set of questions that force the model to utilize at least two modalities (e.g., Graph + Text OR Graph + SMILES) to achieve superior performance compared to models tested only on unimodal data or zero-shot settings.
Sources
- Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
- GPT-4 Technical Report
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- Retrieval-Augmented Generation for Large Language Models: A Survey
- A Survey on LLM-as-a-Judge
- Towards Generalist Prompting for Large Language Models by Mental Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
- GraphAlign: Pretraining One Graph Neural Network on Multiple Graphs via Feature Alignment
- GRAG: Graph Retrieval-Augmented Generation
- GPT-4o System Card
- Mistral 7B
- FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
- From Data to Commonsense Reasoning: The Use of Large Language Models for Explainable AI
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- Synthetic Context Generation for Question Generation
- RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
- GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering