MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

arXiv:2606.06696 · cs.CV, cs.AI · Submitted 2026-06-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models".

Tom: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, but realizing this potential requires robust and fine-grained visual perception across diverse modalities, scales, and contexts.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We've covered a lot of ground today discussing "MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models," and now we get to talk about what the authors are calling it in terms of its overall message.

Jane: That title, "MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models," really tells us that this paper is setting up a very broad field for testing these kinds of AI systems <ref:2606.06695#pg1>.

Lu: I think the real excitement comes from how they’re structuring this benchmark across so many different imaging modalities and biological scales; it opens up so many creative avenues for what these models can actually learn <ref:2606.06695#pg1>.

Meng: From my side, I'm looking at the implications for deployment; if we can get models that perform well here, it means we could start seeing better diagnostic tools in complex clinical settings sooner than we thought <ref:2606.06695#pg4>. We need models that generalize across different types of medical data, not just one specific dataset.

Lalam: I see this as a huge step because it forces the AI to learn how to understand real-world medical data, which will ultimately help improve patient care by making these systems more reliable <ref:2606.06695#pg1>.

Tom: Exactly, and the authors really want us to focus on those core perceptual abilities that they found were lacking in previous evaluations <ref:2606.06695#pg4>.

Jane: They're essentially showing us that we need a much broader range of tests to see how well these vision-language models actually understand medical images in practice <ref:2606.06695#pg2>.

Lu: It suggests that future research should be less about tweaking the language part and more about building better spatial reasoning and perception mechanisms into the AI architecture itself <ref:2606.06695#pg4>.

Meng: And from an engineering standpoint, it means we need to start thinking about how to make these models generalize across different types of medical data, not just one specific dataset <ref:2606.06695#pg4>.

Lalam: If we can solve these issues with better spatial modeling, the impact on healthcare is going to be significant because it lets us trust the AI more in real-world scenarios <ref:2606.06695#pg1>.

Tom: So, MMBU isn't just a benchmark; it’s a tool that clearly shows where current models are falling short in truly understanding complex biomedical information.

Jane: It really puts the pressure on the research community to focus on improving those core visual perception abilities, not just adding more layers to the language part of these models <ref:2606.06695#pg4>.

Conclusion: Tom: So, we've spent some time digging into how MMBU is structured as this massive test for vision and language models in medicine, and now we get to talk about what this whole project is actually called and what it means for us moving forward.

Jane: That title, "MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models," really tells us that the authors are building a comprehensive system to see how these AI systems truly perceive medical images across so many different kinds of data and contexts.

Lu: I think what's fascinating is how they’ve managed to cover thirty-five different modalities and biological scales; it opens up so many creative possibilities for what these models can actually learn from this diverse set of specimens <ref:2606.06695#pg1>.

Meng: From my side, thinking about the actual engineering challenge, this benchmark shows us that we need to stop focusing only on getting high scores on narrow datasets and start building models with genuinely robust spatial reasoning capabilities <ref:2606.06695#pg4>.

Lalam: I agree with that direction for improvement, Meng. The paper points out clear weaknesses in detection and cross-dataset transferability, which gives us a roadmap for what needs to be fixed to improve the AI’s usefulness in real cultural settings <ref:2606.06695#pg4>.

Tom: Exactly, and the main point is that MMBU acts like a really thorough measuring stick because it combines all these different visual tasks into one big assessment for understanding medical data <ref:2606.06695#pg3>.

Jane: They are essentially arguing that we need a much wider variety of tests to see how well these vision and language models actually perform when they encounter real-world medical images, not just in controlled lab environments <ref:2606.06695#pg2>.

Lu: The authors are clearly suggesting that future research should shift its focus away from just improving the language understanding component and move towards building better spatial reasoning mechanisms right into the AI's core architecture <ref:2606.06695#pg4>.

Meng: I see the real impact here as a shift in how we prioritize research, moving from just chasing raw performance numbers to actually understanding the deep visual reasoning capabilities of these AI systems <ref:2606.06695#pg4>.

Lalam: It's about creating AI that can handle the messy reality of medical data, which is where we'll see the most real value in improving how we deliver care to people <ref:2606.06695#pg1>.

Tom: So, this benchmark isn't just a test; it’s a powerful tool that clearly shows us exactly where current vision and language models are falling short when trying to truly grasp the complexity of biomedical information <ref:2606.06695#pg3>. We're going to keep digging into these limitations next.

Stanford University

cs.CV, cs.AI

Submitted: 2026-06-04

Updated: 2026-10-02

Comments: Expanded model results, revised evaluation description, and updated supplementary material

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, but realizing this potential requires robust and fine-grained visual perception across diverse

Key concepts

MMBU Benchmark
This is the largest biomedical vision-language benchmark to date. It covers 35 submodalities across 11 modalities using over 410 datasets. It includes rich structured metadata about image provenance, domain, and acquisition conditions to allow for systematic evaluation of model performance.
Submodalities
These refer to the diverse types of biomedical data being tested, such as different imaging modalities (e.g., X-ray vs. MRI) or biological specimens. The benchmark covers 35 distinct submodalities, ensuring models are tested on a wide range of visual inputs beyond standard datasets.
Closed-to-Open Gap
This measures how much better a model performs when given an answer choice (closed task) compared to when it must generate the answer freely (open task). A large gap suggests the model relies heavily on superficial cues rather than deep, genuine visual understanding of the image.
Object Detection Failure Mode
This highlights a major limitation where models fail to reliably locate objects in images, even in closed-ended tasks. No current VLM surpasses a random baseline score, indicating fundamental weaknesses in spatial reasoning and the ability to pinpoint specific locations within the visual field.

Terminology

Summary

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, but realizing this potential requires robust and fine-grained visual perception across diverse modalities, scales, and contexts. The Massive Multimodal Biomedical Understanding (MMBU) benchmark is introduced as the largest biomedical vision and language benchmark to date to systematically evaluate model performance in this complex domain.

The gist

MMBU is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities.

Introduction and Motivation

Biomedical VLMs must exhibit robust visual perception to capture both fine-grained and coarse-grained features across heterogeneous patient populations and acquisition conditions. Existing evaluations are limited to a small set of commonly used datasets and lack sufficient metadata to stratify performance. Furthermore, training sets underlying current benchmarks often overlap substantially with evaluation data, leading strong performance on established benchmarks to mask deficiencies in visual perception and domain generalization. MMBU addresses this by aiming to provide a more reliable evaluation of biomedical VLMs through broader coverage across visual tasks, imaging modalities, biological specimens, and acquisition conditions.

The MMBU Benchmark Design

MMBU is designed to move beyond narrow domain-specific evaluations by providing comprehensive coverage. Key features of the benchmark include:

  1. Covering 35 submodalities across 11 modalities from 20 specimens and 95 unique regions of interest within those specimens, spanning over 410 datasets.

  2. Each dataset is converted into closed- and open-ended visual question answering (VQA) tasks.

  3. It provides rich, structured annotations for each datapoint, covering image provenance, dataset name, domain, modality, submodality, stain (when applicable), specimen subregion, topic, original task description, context of acquisition institution URL.

  4. It includes both open and closed versions of ungrounded classification and object detection tasks.

Systematic Evaluation Methodology

The evaluation process is structured to allow for systematic analysis across various axes:

  1. The benchmark covers core perception tasks including classification and detection.

  2. Models are selected to include 15 open-weight and 2 frontier VLMs, along with their corresponding base counterparts to assess the impact of biomedical fine-tuning.

  3. Metrics reported include micro-averaged F1-score with 95% confidence intervals obtained via 1,000-iteration bootstrap resampling over datapoints.

  4. For closed-ended VQA, a fully deterministic extraction and scoring protocol is used rather than LLM-as-a-judge evaluation.

Key Findings on Model Performance

The experiments reveal several critical limitations in current models:

  1. Absolute F1 scores remain low overall, with most scores falling below the 0.5 threshold for adequate performance.

  2. The closed-to-open gap is consistently large, with an average delta of 0.26 for classification tasks, indicating models rely heavily on answer-choice cues rather than genuine visual understanding.

  3. Object detection remains a critical failure mode: no VLM surpasses the random baseline (F1 = 0.172) in the closed setting, revealing fundamental limitations in spatial reasoning, as models cannot reliably locate bounding boxes themselves.

  4. Medical adaptation yields limited but measurable gains, and these gains are often domain-selective rather than uniform across all modalities, suggesting that adaptation benefit is strongly tied to modality coverage and data quality.

  5. Models generally perform best on closed classification-related tasks; for instance, Qwen2.5-VL-32B achieves a score of 0.693 on grounded classification from segmentation in the closed setting, but this is an exception for most scores below 0.5.

  6. Cross-Dataset Generalization shows that adaptation gains do not consistently transfer to MMBU; for example, OctoMed improves on both; MedGemma improves on legacy benchmarks but remains flat on MMBU.

Conclusion and Future Directions

MMBU exposes pervasive weaknesses in biomedical VLMs’ perceptual abilities: low ungrounded and grounded classification accuracy, poor detection, and inconsistent cross-dataset generalization. The benchmark highlights clear directions for future work: stronger spatial modeling, improved perception, and adaptation methods that generalize beyond currently established datasets. MMBU provides a representative proxy for measuring progress by consolidating diverse perceptual tasks and rich metadata.

**Table 1: Current landscape of biomedical vision–language benchmarks.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems by leveraging the findings from the MMBU benchmark, along with what those improved systems will be capable of:


  1. Development of a Perception-Aware Layer for Vision-Language Models (VLMs).

  2. Implementation of Fine-Grained Spatial Reasoning Modules (for Object Detection/Segmentation).

  3. Creation of Domain-Adaptive Generalization Frameworks with Modality Sensitivity Control.

  4. Establishment of a Robust, Multi-Level Metadata Extraction and Question Construction Pipeline.

  5. An improved AI system can perform:

  6. Fine-grained spatial localization (e.g., bounding box prediction or segmentation mask generation) on biomedical images, rather than just classification or open-ended answers, achieving an IoU threshold of at least 0.5 (as required by the MMBU Object Detection task).

  7. Accurate identification and classification of specific cellular structures or cell types within complex specimens (e.g., distinguishing between different myeloid lineage cells in a bone marrow smear) by leveraging rich metadata like stain type and specimen subregion, leading to higher accuracy than models relying solely on image content.

  8. An improved AI system can perform:

  9. Systematic performance stratification across diverse biomedical domains (e.g., Oncology, Radiology, Pathology) and imaging modalities (e.g., MRI, Confocal Microscopy).

  10. Robust assessment of domain generalization by quantifying the precise gain or loss in performance when moving from a narrow benchmark to the broad MMBU setting, allowing researchers to identify weak domains where models fail despite high aggregate scores.

  11. An improved AI system can perform:

  12. Adaptive performance tuning based on acquisition context (e.g., scanner type, protocol) by dynamically adjusting its internal weights or prompt strategy to compensate for domain shifts, leading to more reliable performance across different clinical settings (addressing the sensitivity to batch effects noted in the paper).

  13. An improved AI system can perform:

  14. Automated generation of highly specific, clinically grounded visual questions by extracting fine-grained metadata (specimen type, stain, body part/subpart) from raw images and source context to condition a question template, moving beyond generic VQA prompts to generate queries that probe deep visual reasoning.

Abstract

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating 16 open-weight and 5 frontier VLMs in the main comparison, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization. We further define an open-ended MMBU hard set, split into a released public subset and a held-out private subset, to stress-test frontier models.

Sources

Related papers