Vision-language models for chest radiography do not always need the image

arXiv:2606.17710 · cs.CV, cs.AI, cs.CL, cs.LG · Submitted 2026-06-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Vision-language models for chest radiography do not always need the image".

Jane: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: To kick things off, let's look at the title and who wrote this piece. "Vision-language models for chest radiography do not always need the image." It’s a very direct statement about what they found regarding these AI systems.

Jane: The authors are researchers from institutions like TUM University Clinic and RWTH Aachen University, which gives them a strong medical grounding for this kind of work. They are bringing together expertise from both the technical AI side and the clinical radiology side.

Lu: What's compelling about their setup is that they aren't just looking at one type of model; they’re testing this across nine different systems, including specialist medical models and even a text-only version of a large language model.

Meng: That variety is key because it shows that the issue isn't just with one specific architecture, but with the general tendency of these multimodal systems to rely too heavily on linguistic priors when presented with visual data.

Lalam: I think the authors are really showing that high accuracy scores can be misleading because they don't tell you if the model is actually reading the chest radiograph or just guessing based on the words in your prompt.

The paper's summary: Tom: Now, let’s talk about what they actually found in this paper. The core finding is that benchmark accuracy alone isn't a reliable indicator of whether a model uses the image for its medical decisions. They introduce a causal audit to test this dependency directly.

Jane: In simple terms, they are showing that you can have models with very high accuracy purely because they learned to associate certain finding names with certain language descriptions, without actually having the visual evidence in their final answer.

Lu: They test this by running four specific interventions on every question: the original image, swapping it with another patient's scan that has the same label, hiding the radiologist-marked finding region, and hiding a completely irrelevant part of the image.

Meng: That combination of swaps and occlusions is what lets them isolate whether the answer flips when you actually mess with the visual data versus just changing the text context.

Lalam: The key takeaway here is that they found that benchmark accuracy doesn't separate models that are reading the scan from those that are just using linguistic priors, even when some models score quite highly.

The paper's improvements: Tom: The authors suggest a few ways to move forward based on these findings. They propose this causal audit as a way to measure grounding, which they call the "causal triad".

Jane: They are suggesting that we need metrics like CGR, UAR, and IS—the causal grounding rate, unrelated-image answer rate, and irrelevant-mask stability—to sort these systems into behavioral categories.

Lu: The improvements they suggest are very practical for development: first, you need to measure how much of the output is actually governed by the image through those swap interventions. Second, they highlight specific findings like cardiomegaly or pleural effusion that show a consistent signal when tested this way.

Meng: From an engineering perspective, the idea of focusing on the "acquisition geometry," which is whether CGR is higher on posteroanterior versus anteroposterior radiographs, gives us a concrete direction for optimizing image input for better model performance.

Lalam: I think the most important improvement they point toward is measuring accuracy and grounding separately; you can't just look at one number to certify reliability.

Conclusion: Tom: So, to wrap things up, the paper concludes that we have to measure accuracy and image use as two separate things. They stress that reported near-expert accuracies should be read more like a certification of prior-to-dataset alignment rather than proof of radiology itself.

Jane: It’s important to remember that the causal grounding rate is just a conservative lower bound on image use, and some human readers also show modest grounding rates, meaning occluding a single box doesn't necessarily destroy all evidence.

Lu: The whole point of this work is to insist that the assurance we need to have before trusting an AI is seeing evidence that the answer came from the image and not just recited from prior knowledge.

Meng: I think implementing these interventional audits across all deployment pipelines will be a necessary step for any serious clinical application of vision-language models.

Lalam: Ultimately, this paper on "Vision-language models for chest radiography do not always need the image" tells us that we need to move beyond simple leaderboards and start measuring the actual causal dependency of an answer on the visual input.

Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg · Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich · Lab for AI in Medicine, RWTH Aachen University · Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen

cs.CV, cs.AI, cs.CL, cs.LG

Submitted: 2026-06-16

Updated: 2026-10-01

Code: https://github.com/mahshadlotfinia/causa

Project page: https://stanfordmlgroup.github.io/competitions/chexpert

Importance score: 90/100

The gist: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image.

Key concepts

Causal Grounding Rate (CGR)
This metric measures how often a model's correct answer changes when the specific area of interest in the radiograph is hidden or masked. A high CGR suggests the model relies directly on visual information from that region to make its decision, rather than just general text knowledge.
Unrelated-Image Answer Rate (UAR)
This metric checks how often a model's correct answer remains the same even when the entire image is swapped with a different patient's image. A high UAR suggests the model is relying on stable, general linguistic priors rather than specific visual details of the original X-ray.
Irrelevant-Mask Stability (IS)
This assesses how robust a model's answer is when random, non-target areas of the image are occluded. A high IS means the model's conclusion is stable even if irrelevant parts of the image are blocked, indicating it isn't overly dependent on every single visual detail.
Orthogonality of Accuracy and Use
The study found that a model's overall benchmark accuracy does not reliably predict whether it uses the image for its diagnosis. Stronger models that ignore images often outperform those that use them, proving that accuracy alone is an insufficient measure of visual grounding in medical AI.

Terminology

Summary

Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. The central finding of this work is that benchmark accuracy alone cannot distinguish whether a model reads an image or infers from linguistic priors, necessitating an interventional audit to test causal grounding.

The Causal Audit Methodology

The authors introduce a causal audit designed to test whether a correct answer causally depends on the image by intervening on the input. This intervention involves four conditions applied to every question: the original image, a swap to a different patient with the same label, occlusion of the radiologist-marked target region, and occlusion of an irrelevant region. To test for causality, three behavioral metrics are derived from these interventions:

  1. The causal grounding rate (CGR): the fraction of correct-on-original answers that flip when the target region is masked.

  2. The unrelated-image answer rate (UAR): how often a previously correct answer survives an image swap.

  3. The irrelevant-mask stability (IS): how often it survives an irrelevant occlusion.

These metrics are informative only when read together, as they sort the cohort along an axis that accuracy does not reveal. The authors replicate this audit across nine systems, including specialist medical multimodal models, general-purpose multimodal foundation models, frontier closed-source systems, a text-only large language model with no visual encoder (MedGemma-27B-text), and a vision-only linear probe over RAD-DINO image features.

Behavioral Categorization

The audit sorts the nine systems into three distinct behavioral categories based on intervention responses:

  1. Uses image: Defined by "CGR > 0 with a 95% bootstrap interval excluding zero and IS ≥ 90. These systems are characterized by a high CGR but low IS" or similar profiles.

  2. Ignores image: Defined when the model exhibits CGR = 0, UAR = 100, and IS = 100, meaning no edit to the image alters their answers.

  3. Unstable: Defined when "IS < 70," where answers shift under occlusion of any region and CGR cannot be read as localized grounding.

Decoupling Accuracy from Image Use

The core result is that benchmark accuracy and image use are orthogonal. The strongest ignores-image system outscores genuine image users, demonstrating that accuracy alone is not sufficient evidence that a model is doing radiology. For instance, the text-only MedGemma-27B-text at 60.1% accuracy significantly outscores two of the five image users on shared cases. Furthermore, the highest accuracy in the cohort belongs to an image user, Gemma-4-26B at 66.2%, yet the tier just below it mixes categories freely according to behavioral metrics.

Layered Structure of Image Use

The analysis reveals three layers of partiality in how image use is distributed:

  1. Partiality of Governance: how much of a correct output the image actually governs. This is measured by the swap intervention, showing that only 17.9 to 24.7% across the five systems are image-contingent, while the rest are reachable from label-aligned priors.

  2. Finding Specificity: which findings carry the signal. A sparse pattern emerges where only five findings (cardiomegaly, consolidation, edema, pleural effusion, and pneumonia) carry CGR with a Wilson lower bound above zero for every evaluable image user.

  3. Acquisition Geometry: acquisition geometry. CGR is higher on posteroanterior than on anteroposterior radiographs for every image user.

Clinical Implications and Limitations

The findings suggest that accuracy and grounding must be measured separately, and that interventional audits should accompany any clinical-deployment claim rather than relying solely on leaderboards. The study concludes that reported near-expert accuracies should be read as partly certifying prior-to-dataset alignment, not radiology. A key limitation is that the causal grounding rate is a conservative lower bound on image use, and the human reference reader's own grounding rate was modest, indicating that occluding a single box need not remove evidence. Additionally, confidence flags ungrounded answers only when a model uses the image, suggesting confidence-gating is uninformative or anti-calibrated for those models that most need guardrails. The authors emphasize that evidence that a correct answer was read from the image, and not recited from priors, is the assurance that should precede trust.

Key Behavioral Metrics Summary

The three behavioral quantities—CGR, UAR, and IS—are defined as:

(CGR)

the fraction of correct-on-original answers that flip when the target region is masked.

Improvements for AI systems

As a fastidious and diligent researcher, I have thoroughly analyzed this causal audit of Vision-Language Models (VLMs) for chest radiography. The core finding is that benchmark accuracy is an insufficient measure; interventional behavioral audits are necessary to determine if a model actually uses the image or relies on linguistic priors.

Based on these findings, here are specific improvements and capabilities for AI systems:


)1. Implement Mandatory Causal Grounding Audits

Improvement: Integrate the three-pronged causal triad (Target Masking, Swap Intervention, Irrelevant Masking) into the standard evaluation pipeline for medical VLMs.

Capability: Systems must demonstrate a statistically significant Grounding-Specificity Premium (CGR - (1 - IS) > 0). This means a correct answer must flip when the relevant region is occluded (high CGR) and remain stable when an irrelevant region is occluded (high IS).

)2. Establish Behavioral Deployment Tiers

Improvement: Move beyond simple accuracy scores. Deploy models based on their behavioral category determined by the audit: Uses Image, Ignores Image, or Unstable.

Capability: Clinical deployment should be strictly gated. Models in the Ignores Image category (CGR=0, UAR=100, IS=100) are unsuitable for diagnostic tasks requiring visual evidence.

)3. Focus Training on Causal Image Dependency

Improvement: Modify training objectives to explicitly reward causal image use rather than just correlation with finding names. This requires moving beyond standard supervised learning toward causal or attribution-aware training (as suggested by the need for training objectives that explicitly reward causal image use).

Capability: Models will learn to maintain high CGR and IS across various occlusions, making them robust to noise and irrelevant visual distractions, rather than merely memorizing finding-name priors.

)4. Develop Specialized Interpretability Metrics

Improvement: Standard attention maps are insufficient. Implement interventional metrics (CGR, UAR, IS) as primary interpretability indicators for medical VLMs.

Capability: Researchers can quantitatively distinguish between a model that is confidently wrong due to poor image utilization versus one that is correct because it successfully integrated visual evidence.

)5. Refine Confidence Flagging Protocols

Improvement: Adjust confidence flagging logic to be context-aware based on the intervention results.

Capability: Models should only flag ungrounded answers for confidence flags when they demonstrably use the image and fail grounding audits, rather than flagging them universally or based solely on internal probability scores (as high confidence can mask poor evidence utilization).

)6. Contextualize Model Strengths Across Modality/Resolution

Improvement: Develop system-specific performance profiles across different input resolutions (e.g., 224px vs 512px) and prompt styles (default vs terse vs radiologist-framed).

Capability: Clinicians can select the optimal model for the clinical setting—for instance, favoring a model with high CGR at higher resolutions for detailed pathology, or a more robust text-only model for low-resource environments.

)7. Validate Grounding Against Human Expert Performance

Improvement: Use board-certified radiologists as ground truth references to calibrate and validate the interventional metrics (CGR, IS).

Capability: This allows researchers to understand the ceiling of human diagnostic reliability under masking conditions and provides a benchmark against which model grounding can be measured.

Sources

Related papers