Benchmarking Large Language Models for Biomedical Relation Extraction

arXiv:2609.19071 · cs.CL · Submitted 2026-07-23 · Read on arXiv

cs.CL

Submitted: 2026-07-23

Updated: 2026-07-23

License: http://creativecommons.org/licenses/by/4.0/

The gist: Extracting SNP-phenotype associations from biomedical literature is vital but challenging.

Terminology

Abstract

Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen, Mistral), on the SNPPhenA corpus across three tasks: sentence-level, abstract-level, and association strength classification. OpenAI O1 achieved state-of-the-art (SOTA) results using few-shot learning for non-finetuned sentence-level classification (F1 0.89) and established a new SOTA for abstract-level classification (F1 0.82). Association strength classification proved difficult, though fine-tuned Gemini 2.0 Pro performed best (F1 0.60) in the first LLM evaluation of this task. Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models. These findings confirm the power of modern LLMs for genomic knowledge extraction.

Related papers