A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
cs.CL, cs.LG
Submitted: 2025-03-26
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to
Terminology
Abstract
Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data.
Sources
- GPT-4 Technical Report
- MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- {\mu}-Bench: A Vision-Language Benchmark for Microscopy Understanding
- Aviary: training language agents on challenging scientific tasks
- Representation Learning with Contrastive Predictive Coding
- DINOv2: Learning Robust Visual Features without Supervision
- Capabilities of Gemini Models in Medicine
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- HuatuoGPT, towards Taming Language Model to Be a Doctor
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering