MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
cs.CL
Submitted: 2026-03-01
Updated: 2026-09-21
Comments: Technical report, work in progress
License: http://creativecommons.org/licenses/by/4.0/
The gist: Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source
Terminology
Abstract
Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive, precluding the on-premises deployment required for patient privacy and PHI compliance. We introduce MEDGPT-OSS, an open-weight, 20B-parameter generalist vision-language model designed to facilitate open research in clinical AI. Rather than relying on architectural complexity, MEDGPT-OSS pairs the GPT-oss language backbone with a visual front-end via a optimized, three-stage training curriculum. By progressively domain-adapting these modules through rigorous data curation and long-context multimodal alignment, we demonstrate that a 20B model can bridge the capacity gap. It successfully outperforms larger open medical models on out-of-distribution (OOD) multimodal reasoning and complex text-only clinical tasks. By unifying diverse modalities under a single instruction-following interface, MEDGPT-OSS maintains a parameter-efficient footprint fully compatible with commodity GPUs. We release the complete training recipe, open-weight checkpoints, and a rigorous evaluation harness to serve as a verifiable foundation for privacy-preserving, institution-specific clinical AI research.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Qwen Technical Report
- Qwen3-VL Technical Report
- Constitutional AI: Harmlessness from AI Feedback
- Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
- Scaling Laws for Neural Language Models
- Improved Baselines with Visual Instruction Tuning
- GPT-4 Technical Report
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- MedGemma Technical Report
- Kimi K2.5: Visual Agentic Intelligence
- Towards Generalist Biomedical AI
- Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
- DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
- MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
- TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering