From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "From Literature to Hypotheses".
Tom: CoDHy, an interactive AI co-scientist system,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We’ve covered the general idea, but let’s look closer at who put this together. The authors are Raneen Younis, Suvinava Basak, Lukas Chavez, and Zahra Ahmadi from institutions like the Peter L. Reichertz Institute for Medical Informatics and the Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine.
Jane: It’s interesting to see a collaboration spanning different medical informatics and AI research centers; that suggests they're bringing together different expertise to solve this problem from multiple angles.
Lu: The team’s background seems perfectly suited for this task, blending deep knowledge of biomedical databases with advanced causal methods and artificial intelligence techniques. That kind of cross-disciplinary strength is what makes the system's architecture so compelling.
Meng: I wonder if having authors from both the medical informatics side and the core AI research side helps them balance the need for scientific accuracy with cutting-edge machine learning capabilities in a very practical way.
Lalam: I think their diverse expertise really reflects how complex this problem is; it isn't just one type of challenge, it requires knowledge spanning data structuring, natural language understanding, and causal inference all at once.
The paper's summary: Tom: Now that we know who’s behind the work, let’s talk about what they actually did in "From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation." Essentially, they built an interactive system called CoDHy that takes a focus biomarker and a cancer type, lets you specify how much literature to look at, and then it builds a knowledge graph from everything.
Jane: That knowledge graph is the core engine; it’s not just a static database, but it dynamically pulls in unstructured information from PubMed to create this task-specific structure that guides the whole process. It’s like they are creating a custom map for that specific research question.
Lu: The summary highlights how CoDHy integrates several steps: building the graph, learning embeddings on it, then having agents generate and validate hypotheses using both the graph structure and those learned mathematical representations. It’s a multi-layered approach to scientific discovery.
Meng: The description emphasizes that the system doesn't just spit out random ideas; it uses this combination of structured evidence and learned embeddings to propose drug combinations, which is a significant step toward generating hypotheses with some degree of contextual grounding.
Lalam: What stands out from the summary is that they aren't stopping at generation; they have validation and ranking agents that check the hypotheses for plausibility and novelty before presenting them back to the user. That iterative refinement capability sounds really powerful for real research workflows.
The paper's improvements: Tom: The paper also points out how they improved this concept by suggesting specific technical enhancements, moving beyond just a basic setup. They propose things like using modular pipelines so you can swap out components, and implementing a dynamic cache that intelligently clears old data based on what you’re currently investigating.
Jane: That idea of the intelligent invalidation mechanism for the knowledge graph cache is smart because it means researchers don't have to wait for a full system rebuild every time they tweak their input parameters; it makes the iteration process much smoother.
Lu: And I think their suggestion to incorporate relation-aware embeddings like RotatE or ComplEx instead of just Node2Vec shows they are thinking about capturing the complex relationships between entities more accurately, which is crucial when you're dealing with nuanced biomedical interactions.
Meng: From a practical perspective, integrating those relation-aware models means that the system can model the multi-relational semantics of drugs and biomarkers better than a purely structural approach would allow, which should lead to more meaningful connections in the hypothesis generation phase.
Lalam: I agree with Lu; improving how the embeddings capture semantics directly affects how good those hypotheses are before they even get validated by the other agents; it’s about making the latent signal richer from the start.
Conclusion: Tom: So, wrapping up this discussion on "From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation," we see a system that successfully merges structured data, literature mining, and agent reasoning into one interactive framework. The paper shows how this structure helps researchers get evidence-grounded hypotheses in oncology scenarios.
Jane: It’s clear the main implication is moving from a manual search and connection process to an automated discovery process where the AI acts as a true co-scientist, helping to propose testable drug combinations with explicit rationales.
Lu: The shift toward this co-scientist model suggests that future research won't just be about analyzing existing literature, but about actively proposing novel avenues based on what the system can infer from that literature and structured data.
Meng: For practical impact, this means researchers can spend less time manually synthesizing connections and more time focusing their experimental resources on the most promising, evidence-supported drug combinations generated by this kind of system.
Lalam: I really think the cultural impact here is shifting how we view AI's role in science; it’s moving from a tool that just summarizes to a partner that can actively generate and critique scientific ideas with both direct evidence and inferred context.
Peter L. Reichertz Institute for Medical Informatics (PLRI) · Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine (CAIMed) · Sanford Burnham Prebys Medical Discovery Institute (SBP) Medical Discovery Institute
cs.CL
Submitted: 2026-02-28
Updated: 2026-10-01
Code: https://github.com/baksho/CoDHy
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 88/100
The gist: CoDHy, an interactive AI co-scientist system, addresses the challenge of systematically connecting biomarker mechanisms to actionable drug combination hypotheses in cancer research by integrating
Key concepts
- Knowledge Graph Construction
- This involves building a unified map of biomedical information by combining structured data (like databases) and unstructured text from PubMed. It extracts entities like biomarkers and drugs, and relations like 'targeting' from the text, storing them in a graph database for systematic reasoning.
- Node2Vec Embeddings
- These are numerical representations of every entity (node) in the knowledge graph. Node2Vec captures how nodes are connected by modeling random walks across the graph structure. This allows the system to understand structural relationships and calculate quantitative similarity between different biomedical concepts, even if they aren't directly linked.
- Graph-RAG Approach
- This is a method where an AI agent generates hypotheses by first finding specific, explicit connections in the knowledge graph related to a biomarker. It then enhances this information by looking at similar nodes in the embedding space (latent relationships), allowing it to generate hypotheses based on both direct evidence and implied connections.
- Graph Evidence Score
- This is the final metric used to rank drug combination hypotheses. It is a weighted score combining three factors: how many direct edges support the hypothesis, how strong the similarity between concepts is in the embeddings, and how much evidence covers the topic. This composite score determines which hypotheses are most promising.
Terminology
Summary
CoDHy, an interactive AI co-scientist system, addresses the challenge of systematically connecting biomarker mechanisms to actionable drug combination hypotheses in cancer research by integrating structured biomedical databases and unstructured literature evidence into a task-specific knowledge graph. This system generates, validates, and ranks candidate drug combinations using graph-based reasoning and multi-agent systems while maintaining a human-in-the-loop interaction for transparent exploration.
System Overview
The CoDHy system is designed as an interactive, AI co-scientific tool that combines a web interface with a modular backend pipeline for biomarker-guided drug combination hypothesis generation. The core function involves constructing a task-specific knowledge graph from structured resources and unstructured PubMed literature, followed by graph-based reasoning and validation. The architecture integrates (i) task-specific knowledge graph construction, (ii) graph-based representation learning, (iii) multi-agent hypothesis generation and validation, and (iv) human-in-the-loop interaction within a unified framework. Users configure the scientific context—specifying a focus biomarker,
cancer type or disease context,
and the number of PubMed abstracts to retrieve
—to initiate the process.
Knowledge Graph Construction
The backend pipeline begins by constructing a unified biomedical knowledge graph that integrates both structured and unstructured evidence sources. Curated biomedical resources are ingested as typed entities and relations via APIs, while unstructured evidence is collected from PubMed based on user-controlled queries. Biomedical entities (e.g., biomarkers, drugs, pathways
) and relations (e.g., targeting, activation
) are extracted using spaCy-based NLP pipelines from the retrieved abstracts. To ensure graph consistency against noisy relation extraction, extracted relational statements are encoded using sentence-transformers and mapped to the most semantically similar existing relation type via cosine similarity. The resulting graph is stored in Neo4j AuraDB for efficient querying and structured reasoning, with a graph cache maintained for previously constructed sets of biomarkers and cancer types to reuse across subsequent cycles.
Graph Embedding Generation
To enable similarity-based reasoning over the constructed knowledge graph, node embeddings are computed using Node2Vec. This method captures structural proximity through biased random walks, modeling both local neighborhood structure and global connectivity patterns, providing efficient structural representations suitable for dynamically constructed graphs. These embeddings allow the system to compute quantitative similarity between biomedical entities,
supporting inference beyond explicitly encoded edges. While Node2Vec is adopted for its computational efficiency, the framework is modular and can incorporate relation-aware models such as RotatE or ComplEx to better capture multi-relational semantics.
Hypothesis Generation Agent
Hypothesis generation is performed by a dedicated agent operating over both the symbolic knowledge graph and its learned embeddings using a hybrid Graph Retrieval-Augmented Generation (Graph-RAG) approach.
The process involves two steps: first, the agent retrieves a localized discovery subgraph of explicit biomedical interactions connected to the focus biomarker,
representing direct evidence. Second, this context is augmented with implicit signals by identifying the most similar nodes to the input biomarker in the embedding space (via cosine similarity),
enabling generation based on latent relationships. All hypotheses are produced in a standardized output format, allowing for downstream validation and ranking.
Hypothesis Validation Agent and Ranking
Each generated hypothesis is evaluated by a validation agent that assesses novelty, plausibility, and feasibility using an LLM-based reasoning process.
The agent explicitly distinguishes between hypotheses supported by direct graph evidence
versus those inferred via embedding similarity.
For all candidates, the validation agent performs targeted PubMed searches to support or challenge the hypothesis and assigns a qualitative verdict (proceed, caution, or reject
) along with a safety assessment score. In the final stage, a ranking agent computes a graph evidence score,
which is a weighted aggregation of (i) direct edge support count, (ii) embedding similarity strength, and (iii) evidence coverage indicators,
combined with an LLM-derived safety score to produce the final composite score for ranking.
Evaluation and Findings
Controlled case studies demonstrate CoDHy’s ability to generate evidence-grounded and context-aware hypotheses in realistic oncology scenarios.
Comparative experiments show that the Full CoDHy system achieves the highest exact novelty (35.71%) while maintaining high combination diversity (0.89) and competitive evidence coverage. Furthermore, when compared to an LLM-only baseline, removing embedding-based inference (No-Node2Vec
) increases exact novelty (28.57%) while maintaining strong evidence grounding (0.72), suggesting that integrating structured knowledge graphs with embedding-based inference enables a better balance between novelty, evidence support, and hypothesis diversity. The system successfully transitions from a search-retrieval paradigm toward a discovery-oriented paradigm,
as evidenced by the Full CoDHy system achieving an MRR of 0.74 compared to 0.93 for the LLM-Only baseline.
Improvements for AI systems
Here are specific, actionable improvements for existing or future AI systems, based on the CoDHy framework presented in the paper:
-
Use a modular pipeline approach that decouples knowledge graph construction from embedding generation and agent reasoning to allow independent updates to each component (e.g., swapping Node2Vec for RotatE without rebuilding the entire system).
-
Implement a dynamic
knowledge graph cache
with an intelligent invalidation mechanism based on changes in the input biomarker or cancer type, rather than relying solely on user configuration changes, to minimize redundant processing cycles during iterative refinement. -
Integrate relation-aware embeddings (like RotatE or ComplEx) as a primary embedding generation method for complex biomedical reasoning tasks to better capture multi-relational semantics compared to structural methods like Node2Vec alone.
-
Enhance the Hypothesis Generation Agent by incorporating explicit
uncertainty quantification
metrics derived from both the graph evidence score and the embedding similarity strength, allowing downstream agents to weigh exploratory (embedding) versus grounded (graph) hypotheses more intelligently before validation. -
Develop a sophisticated safety assessment module that goes beyond simple toxicity scoring by performing automated cross-referencing against SIDER/DrugBank data for specific combination side effects, flagging potential contraindications earlier in the process.
-
Introduce a
Human-in-the-Loop Feedback
mechanism where the researcher's manual interaction during iterative refinement (e.g., rejecting a hypothesis) is immediately used to fine-tune the prompt/reasoning parameters of the LLM validation agent for subsequent runs, creating a personalized feedback loop for hypothesis generation. -
Expand the evaluation beyond just novelty and coverage by incorporating metrics that specifically measure
mechanistic diversity
(e.g., clustering hypotheses by underlying biological pathway) to ensure generated hypotheses cover a broad spectrum of potential therapeutic strategies rather than focusing on narrow, repetitive connections. -
Develop an automated system for generating synthetic, negative evidence (counter-hypotheses) based on contradictory literature retrieved during the validation phase to proactively test the robustness of a candidate drug combination before committing resources to experimental testing.
Abstract
The rapid growth of biomedical evidence makes it difficult to translate biomarker mechanisms into actionable drug combination hypotheses. We present CoDHy, an interactive AI co-scientist for biomarker-guided hypothesis generation in oncology. CoDHy constructs task-specific knowledge graphs from curated databases and biomedical literature, then combines graph embeddings with agent-based reasoning to generate, validate, and rank evidence-grounded drug combinations. Through a web interface, researchers specify the biomarker, cancer context, and literature scope; inspect supporting evidence and intermediate results; and iteratively refine the generated hypotheses. The demonstration presents CoDHy's end-to-end workflow and shows how researchers can interactively explore and compare mechanistically supported drug combinations while remaining in control of hypothesis prioritization.
Sources
- AI4Research: A Survey of Artificial Intelligence for Scientific Research
- TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
- Scientific Hypothesis Generation and Validation: Methods, Datasets, and Future Directions
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space
- BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation
- MIRAI: Evaluating LLM Agents for Event Forecasting
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering