Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510.15042 · cs.CV, cs.LG · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Comprehensive language-image pre-training for 3D medical image understanding".

Jane: The paper was written by Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma et al. from Microsoft and German Cancer Research Center (DKFZ) and University of Cambridge and Cambridge University Hospitals NHS Foundation Trust and Heidelberg University Hospital and Mayo Clinic.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we’re cracking open a paper that’s got a mouthful of a title: “Comprehensive language–image pre-training for three dee medical image understanding.” Jane, I’ll be honest, when I first saw “three dee medical image understanding,” I thought, okay, we’ve seen CLIP for photos, but this is a whole different beast.

Jane: It really is, Tom. And that title is doing a lot of work. “Comprehensive” is the key word here. It’s not just about matching a picture to a caption. It’s about building an AI that can actually read a CT scan and understand what it’s seeing, the way a radiologist would.

Tom: Right, and the team behind this is a who’s who. You’ve got folks from Microsoft Research, the German Cancer Research Center, Cambridge, and Mayo Clinic. That’s a serious collaboration.

Jane: Absolutely. And they’re tackling a problem that’s been a real bottleneck. For 2D images like X-rays, we have tons of paired data—image and report. But for three dee CT scans, that paired data is scarce. It’s like trying to teach someone a language with only a few textbooks.

Tom: So they’re not just making a bigger model, they’re making a smarter training recipe. The title says “comprehensive” because they’re pulling in every trick they can to squeeze value out of the data they have.

Jane: Exactly. And that’s what we’re going to dig into today. How they’re combining different learning objectives to make a model that’s not just good at one thing, but can handle classification, retrieval, and even generating reports.

Tom: And the implications are huge. Imagine a tool that can help a radiologist in a busy hospital, flagging potential issues in a scan, or finding similar past cases to compare against.

Jane: Or even generating a draft report to save time. That’s the promise here. But before we get to the results, let’s talk about the problem they’re solving. Why is three dee so much harder than 2D?

Tom: Well, for starters, a CT scan isn’t one picture. It’s hundreds of slices stacked together. That’s a massive amount of data, and it takes a lot of computing power to process.

Jane: And the reports that come with them are long and detailed, not like a simple caption. So you have this huge gap between the complexity of the image and the complexity of the text.

Tom: Plus, the data is locked away in hospitals. It’s not like scraping the internet for cat pictures. So the authors had to be clever about how they used what was available.

Jane: And that’s the core of the paper. They’ve found a way to use both paired image-report data and a much larger pool of images that have no reports at all.

Tom: So, they’re getting more out of less. That’s the kind of efficiency that could really move the field forward.

Jane: It could. And in the next segment, we’ll break down exactly how they did it. Stick around.

Summary: Jane: So, Tom, we’ve set the stage. The paper is “Comprehensive language–image pre-training for three dee medical image understanding,” and the challenge is the data. Let’s talk about what they actually built.

Tom: Right. They call their model family COLIPRI. And the core idea is to take the CLIP paradigm—that’s the one that aligns images with text—and adapt it for three dee medical scans.

Jane: But they didn’t just copy-paste CLIP. They had to make some serious adjustments. For one thing, the reports are super long. The average findings section is over two hundred forty tokens long. That’s a lot of text to align with a three dee volume.

Tom: And they found that if you train on those long reports, the model gets confused when you ask it a short question like “lung nodule present?”. There’s a distribution shift.

Jane: So they did something clever. They used a large language model to restructure the reports into eight clinical subsections, like “Lungs and Airways” or “Cardiovascular Structures.” Then they broke them down into short, positive and negative findings.

Tom: That’s the preprocessing. But the real magic is in the training objectives. They combined four different losses. The first is the standard CLIP contrastive loss.

Jane: The second is their new “Opposite Sentence Loss.” This is where they take a positive finding, like “lung nodule,” and create a negative one, “no lung nodule.” They train the model to know which one matches the image.

Tom: That’s a great idea for zero-shot classification. It teaches the model to understand the presence or absence of a finding, which is exactly what you need for a quick clinical query.

Jane: And the third objective is report generation. They don’t just want the model to match an image to a report; they want it to be able to generate the report from the image. That forces the vision encoder to capture all the details, not just the most obvious ones.

Tom: And finally, they add a vision-only objective, a masked autoencoder. This is where they take an image, mask out parts of it, and ask the model to fill in the blanks.

Jane: And this is the key to using all that unpaired image data. They can train on thousands of CT scans that don’t have reports, which helps the model learn better spatial features.

Tom: So they’ve got this multi-pronged approach. They’re not just relying on one signal. They’re combining global alignment with local detail and language understanding.

Jane: And the results, which we’ll get into, show that this combination is more powerful than any single approach. The model they call COLIPRI-CRM, which uses all four objectives, is their best performer.

Tom: It’s like they’re giving the AI a complete education. It learns to see, to read, and to describe. And that’s what makes it so versatile.

Jane: Exactly. And in the next segment, we’ll look at the specific improvements and the numbers that back this up.

Tom: Let’s take a short break, but we’ll be right back with the results.

Improvements: Tom: We’re back with “Comprehensive language–image pre-training for three dee medical image understanding.” Jane, we’ve talked about the recipe. Now let’s talk about the proof. Did it actually work?

Jane: It did, and the improvements are substantial. Let’s start with zero-shot classification. This is where you ask the model to identify an abnormality without any fine-tuning. Their best model, COLIPRI-CRM, hit an AUROC of around eighty percent on the CT-RATE dataset. That’s a big jump over the previous state-of-the-art, which was around seventy-seven percent.

Tom: And that’s with short prompts, like “lung nodule present.” That’s a huge deal for practical use. You don’t need to write a paragraph to get a useful answer.

Jane: Right. And it’s all thanks to that Opposite Sentence Loss we talked about. It closes the gap between the long reports used in training and the short queries used in real life.

Tom: But it’s not just about classification. What about report generation? That’s where they really shine.

Jane: Oh, absolutely. When they fine-tune a language model on top of their vision encoder to generate reports, the clinical accuracy is much higher. They use a metric called RadFact, which checks if the generated statements are factually correct.

Jane: Their model improved the F1 score for positive findings by about nine points over the strongest baseline. That means the generated reports are far more likely to correctly identify abnormalities that are present.

Tom: So it’s not just generating grammatically correct text; it’s generating clinically accurate text. That’s a massive step.

Jane: And then there’s retrieval. If you have a report, can you find the matching image? Their model is dramatically better at this, with a recall at five of over thirty-five percent, compared to under three percent for the older CT-CLIP model.

Tom: That’s more than a tenfold improvement. That kind of capability could be used to find similar cases for a radiologist to review.

Jane: And finally, they tested on semantic segmentation, which is the pixel-level task of outlining organs or tumors. Here, the inclusion of the masked autoencoder objective is crucial. It gives the model the fine-grained spatial understanding needed for this task.

Tom: So, each objective has its role. The contrastive loss for global understanding, the OSL for zero-shot, the report generation for clinical detail, and the MAE for dense tasks.

Jane: Exactly. It’s a complete package. But it’s not perfect. They note that the performance, while state-of-the-art, is still below what’s needed for autonomous clinical use.

Tom: Right, it’s a tool to assist, not replace, the radiologist. But the potential is clear. And this brings us to the bigger picture.

Jane: Let’s hear what our other hosts think about the implications. Lu, Meng, what’s your take?

Lu: I think the most exciting part is the training paradigm. The idea of combining multiple objectives to overcome data scarcity is going to be a blueprint for other medical imaging domains, not just CT.

Meng: From an engineering standpoint, the fact that they release the model weights is huge. It means we can start building applications on top of this immediately, without having to train a model from scratch.

Tom: And that’s the kind of impact that moves the field forward. Let’s wrap this up in our final segment.

Conclusion: Tom: We’re in the home stretch now, talking about “Comprehensive language–image pre-training for three dee medical image understanding.” Jane, it’s been a fascinating discussion.

Jane: It really has, Tom. To sum it up, this paper isn’t just about building a better model. It’s about a smarter way to train it. They’ve shown that by combining contrastive learning, a novel opposite sentence loss, report generation, and masked autoencoding, you can create a single encoder that’s a jack-of-all-trades.

Tom: And a master of most, too. It excels at classification, retrieval, and even generating clinically accurate reports. The key is that they’ve managed to leverage both paired and unpaired data, which is a huge advantage given the data constraints in medicine.

Jane: The impact is clear. This could lead to better decision support tools for radiologists, more efficient workflows, and potentially even help in underserved areas where expert radiologists are scarce.

Lu: I’d add that the methodology will inspire more research into multi-objective pre-training for other three dee modalities like MRI. The principles are transferable.

Meng: And the fact that it’s open-source means we’ll see a wave of innovation built on this foundation. It lowers the barrier to entry for startups and research labs alike.

Lalam: From a cultural perspective, this is a step toward democratizing medical expertise. It doesn’t replace the doctor, but it gives every clinician a powerful second opinion, making high-quality diagnostic support more accessible globally.

Tom: That’s a powerful vision. So, we’ll say goodbye to COLIPRI. It’s a significant leap forward for three dee medical imaging.

Jane: It is. And we’re excited to see what the community builds with it. Thanks for listening, everyone. We’ll be back soon with another paper.

Tom: Until next time, keep exploring.

Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando Pérez-García

Microsoft · German Cancer Research Center (DKFZ) · University of Cambridge · Cambridge University Hospitals NHS Foundation Trust · Heidelberg University Hospital · Mayo Clinic

cs.CV, cs.LG

Submitted: 2026-08-15

Updated: 2026-08-18

Code: https://github.com/MIC-DKFZ/nnssl

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

Key concepts

COLIPRI
The model family developed in the paper. It adapts the CLIP paradigm for 3D medical scans by combining multiple learning objectives to handle classification, retrieval, and report generation.
Opposite Sentence Loss (OSL)
A new training objective where the model is trained to distinguish between a positive finding sentence (e.g., 'lung nodule') and a negative one ('no lung nodule'). This helps the model understand the presence or absence of specific findings for zero-shot classification.
Report Generation
An objective where the model is trained to generate a clinical report from an image. This forces the vision encoder to capture detailed spatial information, leading to clinically accurate text generation.
Masked Autoencoder (MAE)
A vision-only objective where parts of an image are masked, and the model must fill in the missing details. This helps the model learn fine-grained spatial understanding, which is crucial for tasks like semantic segmentation.

Terminology

Summary

Summary

This paper introduces the Comprehensive Language–Image Pre-training (COLIPRI) encoder family for 3D medical image understanding, specifically focusing on chest CT scans. The authors address the challenges of adapting vision–language pre-training (VLE) to the 3D medical domain, which are primarily the lack of large paired image–text datasets and the high computational cost of processing 3D volumes.

The paper's contributions are fourfold. First, the authors investigate key design choices of the CLIP training paradigm for 3D medical imaging and propose a novel Opposite Sentence Loss (OSL) to enable zero-shot classification with short queries. Second, they introduce a radiology report generation (RRG) objective akin to CapPa to increase supervision from medical reports. Third, they demonstrate that combining a vision-only self-supervised objective (a masked autoencoder, MAE) with the CLIP objective allows the inclusion of image-only data and enables more localised supervision for dense downstream tasks. Fourth, they comprehensively evaluate the resulting models on zero-shot classification, classification probes, report generation, retrieval, and semantic segmentation.

The COLIPRI framework combines four learning objectives: (a) contrastive image–report alignment (CLIP), (b) the novel Opposite Sentence Loss (OSL), (c) radiology report generation (RRG), and (d) masked autoencoders (MAEs). The models are developed in stages, each adding an objective to assess the effects of each building block. The intermediate models are COLIPRI-C (CLIP + OSL), COLIPRI-CR (COLIPRI-C + RRG), COLIPRI-CM (COLIPRI-C + MAE), and the final COLIPRI-CRM (all objectives combined).

Key methodological details include: using a Primus-M transformer as the vision encoder and CXR-BERT as the text encoder; preprocessing reports with a large language model (LLM) to structure them into eight clinical subsections and to create concise positive and negative findings sentences; and using input sizes of 160×160×160 voxels at 2 mm isotropic spacing for vision-language training. The OSL is designed to reduce the domain shift between long training reports and short zero-shot prompts by training the model to differentiate between semantically opposing statements (e.g., abnormality present vs. no abnormality present). The RRG objective uses a lightweight EVA-02 transformer decoder with parallel captioning. The MAE objective is trained on both paired (CT-RATE) and unpaired (NLST) data, using high-resolution (1 mm) and default-resolution (2 mm) sub-crops.

The models are pre-trained on CT-RATE (for vision-language and vision-only objectives) and NLST (for vision-only objective). Evaluation is performed on multiple tasks. For classification probes, COLIPRI-CM and COLIPRI-CRM achieve the best results, exceeding all baselines on both CT-RATE and RAD-ChestCT datasets. For example, COLIPRI-CRM increases AUPRC by 6 points and AUROC by 3.5 points over the best baseline (Merlin). For zero-shot classification, COLIPRI-CM and COLIPRI-CRM outperform all baselines on CT-RATE, and COLIPRI-CRM generalises better on RAD-ChestCT. The OSL is shown to close the gap between native long-form prompting and short-form prompting. For report generation, COLIPRI-CRM improves RadBERT MacroF1 by 23 points and RadFact-CT (+)/F1 by 9 points over the strongest baseline, indicating more clinically accurate reports. For report-to-image retrieval, COLIPRI-CRM achieves an R@5 of 35.1% compared to 2.9% for CT-CLIP. For semantic segmentation, COLIPRI-CM and COLIPRI-CRM perform on par or better than the state-of-the-art MAE pre-training method, and COLIPRI-CRM exceeds the pure MAE baseline when fine-tuned for 250k steps.

The authors conclude that COLIPRI achieves state-of-the-art performance across all standard tasks, including classification, retrieval, segmentation, and report generation. They note limitations, including that the RRG objective yields only slight improvements and that clinical performance metrics remain below typical clinical thresholds, suggesting that training on larger and more diverse datasets would further enhance performance. The model is available at https://huggingface.co/microsoft/colipri.

Improvements for AI systems

Based on the scientific paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvements:

  • Implement the COLIPRI-CRM architecture combining four objectives: contrastive image–report alignment, Opposite Sentence Loss (OSL), radiology report generation (RRG), and masked autoencoder (MAE).

  • Use Primus-M transformer as vision encoder and CXR-BERT as text encoder with multi-head attention pooling.

  • Pre-train on CT-RATE (paired) and NLST (image-only) datasets with alternating vision–language and vision-only objectives.

  • Include Sentence Shuffle and Short Sentence augmentations to reduce domain shift between long reports and short zero-shot prompts.

What the improved system can do:

  • Zero-shot classification of 18 chest CT abnormalities with short prompts (e.g., Lung nodule present) achieving 80% AUROC on CT-RATE, outperforming prior state-of-the-art by 3–5 points.

  • Report-to-image retrieval with R@5 of 35.1% vs. 2.9% for CT-CLIP, enabling radiologists to find similar cases from text queries.

  • Semantic segmentation after fine-tuning, achieving Dice scores on par with or exceeding MAE pre-training on LiTS, Lung, HVS, and KiTS23 datasets.

  • Report generation with 23-point improvement in RadBERT MacroF1 over baselines, producing clinically more accurate findings.


Overall, the improved AI system can:

  • Serve as a general-purpose 3D medical image encoder for chest CT, supporting classification, retrieval, segmentation, and report generation.

  • Assist radiologists by retrieving similar cases, predicting abnormality likelihoods, and generating draft reports with higher clinical accuracy.

  • Operate with short, natural-language queries in real-time clinical workflows.

  • Generalize across datasets (CT-RATE, RAD-ChestCT) and tasks, setting a new state-of-the-art benchmark.

Sources

Related papers