Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
cs.CL, cs.AI, cs.CV
Submitted: 2025-06-08
Updated: 2026-09-21
Comments: Accepted by TPAMI. Our webpage is https://alibaba-damo-academy.github.io/lingshu. Models and training data are available at https://huggingface.co/lingshu-medical-mllm
Code: https://github.com/xuehuachunsheng/DupImageDetection
Project page: https://alibaba-damo-academy.github.io/lingshu
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- OpenAI ChatGPT interprets Radiological Images: GPT-4 as a Medical Doctor for a Fast Check-Up
- Qwen2.5-VL Technical Report
- CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- The 2024 Brain Tumor Segmentation (BraTS) Challenge: Glioma Segmentation on Post-treatment MRI
- Gemini: A Family of Highly Capable Multimodal Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- The Llama 3 Herd of Models
- PanNuke Dataset Extension, Insights and Baselines
- Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
- GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- MAIRA-1: A specialised large multimodal model for radiology report generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering