Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions".
Jane: The paper was written by Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng et al. from University College Dublin and The Third Affiliated Hospital of Southern Medical University and Dublin City University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everybody. Today we're looking at a paper that's got a real mouthful of a title: "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions."
Jane: And Tom, that title is actually doing a lot of work. Let me break it down. "Vision-language models" are those AI systems that can look at an image and then answer questions about it in text. So here, they're looking at chest X-rays.
Tom: Right, and the models are saying things like "this patient has fluid in their lungs" or "no sign of collapse." But the key word in that title is "auditing." These models are being checked, like a financial audit, to see if their answers actually hold up.
Jane: And the "across institutions" part is the twist. A model might be great at one hospital, but when you move it to a different hospital, with different patients, different equipment, different radiologists writing the reports, does it still get things right?
Tom: So it's not just "is the AI smart?" It's "is the AI trustworthy when I bring it into my building?" And the paper found some pretty wild stuff. I mean, one model literally stopped saying anything was wrong at one hospital. It just said "no findings" for everything.
Jane: That's the kind of thing that would terrify a radiologist. You think you have a tool that's helping you, and it's actually just nodding along. But the paper is careful to say they're measuring agreement with the hospital's own labels, not whether the patient is actually sick.
Tom: Right, that's a crucial distinction. They're not saying "this model is clinically wrong." They're saying "this model disagrees with your hospital's established readings, and the disagreement pattern changes depending on where you are."
Jane: And that's why the title matters. It's not "are these models good?" It's "can you measure how much to trust them at your specific site?" That's a much more practical question for a hospital administrator.
Tom: Exactly. And the answer they found is, honestly, a bit uncomfortable. You can't just borrow trust from another hospital. You have to measure it yourself, with your own data.
Jane: So stick around, because we're going to get into how they actually tried to do that measurement, and whether their clever estimation tricks actually worked. Because spoiler alert, some of them didn't.
Tom: That's the hook, folks. The models are powerful, but trust doesn't transfer. Next up, we're going to dig into what the paper actually did and what they found.
Summary: Tom: So Jane, we've set the stage. The paper is "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." Now let's talk about what they actually did. This is a massive study.
Jane: Massive is the word. They took three different vision-language models, ran them on three different chest X-ray datasets, and looked at six different medical findings. That's like three different brands of X-ray glasses, tested on three different hospitals' patient populations.
Tom: And they didn't just ask the models once. They asked them two different ways. One way was a narrative prompt, like "describe what you see." The other was a binary prompt, like "is there edema? Yes or no?" And the way you ask matters. A lot.
Jane: That's one of the most striking findings, actually. The same model, on the same image, would give a different answer depending on how the question was phrased. We're talking about a kappa score of zero point two, which means the two protocols barely agreed with each other.
Tom: And it gets worse. One of the models, LLaVA-Med, under the narrative prompt, never said a single finding was present. Not once, in over sixty thousand judgments. But under the binary prompt, it said every single finding was present. It went from "nothing is wrong" to "everything is wrong" just by changing the question format.
Jane: That's a complete breakdown. And it tells you that these models aren't really "reading" the X-ray the way a human does. They're responding to the prompt structure in ways that are hard to predict.
Tom: But here's the thing, Jane. The paper isn't just documenting these failures. They're trying to solve a practical problem. If I'm a hospital, and I want to use one of these models, how do I know if I can trust it at my site?
Jane: And their answer is, you need to test it yourself. You need to take a small sample of your own X-rays, get your own radiologists to label them, and see how the model does against those labels. They call this a "target label budget."
Tom: So you spend maybe a hundred or two hundred X-rays, you label them yourself, and then you try to estimate the model's agreement rate at your hospital. And the paper compares ten different ways of doing that estimation.
Jane: And that's where the results get really interesting, because the clever methods didn't always win. We'll get into that in the next segment, but the short version is, sometimes the simple approach is the best approach.
Tom: And sometimes, no approach is good enough. The uncertainty intervals they built, the ones that are supposed to tell you "we're ninety-five percent sure the agreement is between here and here," they only covered the truth eighty-seven percent of the time.
Jane: So even when you do the work, even when you label your own data, the estimates can still be off. That's a sobering message for anyone who wants to deploy these models.
Tom: So the summary is: the models are inconsistent, the measurement is hard, and the trust has to be earned locally. Next, we're going to talk about the improvements the paper suggests, and whether they actually hold up.
Improvements: Tom: Welcome back. We're still on "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." And Jane, I want to dig into the "improvements" part of the paper. Because they don't just complain, they actually propose a method.
Jane: Right, they lay out a whole procedure. It's in the paper as Algorithm one. It's basically a step-by-step guide for a hospital to follow when they want to adopt one of these models.
Tom: And the first step is the most important one, I think. They call it a "fit-for-purpose gate." Before you even bother with fancy estimators, you check whether any estimator beats just using a single average number for everything.
Jane: And that's a really practical move. Because if the model is so bad that the best you can do is a constant "it's right about ninety percent of the time," then you don't need a complex statistical model. You just need that number.
Tom: And the paper found that this gate actually matters. In four out of two hundred and forty decisions, the gate correctly said "stop, just use the constant." That's not a huge number, but it shows the gate is doing something.
Jane: But here's where the improvements get messy. They designed a smart selection procedure. It looks at the available data, it ranks the different estimators, and it picks the best one for your site. And that sounds great in theory.
Tom: But in practice, it failed. The adaptive selection procedure had a Brier score of zero point one zero eight three. But if you just always used one simple estimator, the Beta-Binomial one, you got zero point zero eight five three. The adaptive procedure was worse than just picking a default and sticking with it.
Jane: And that's a really important negative result. It means that trying to be clever about which estimator to use, at least with the data available, actually hurt. The selection process itself introduced error.
Tom: Right, and the paper is honest about this. They say "adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives." That's a direct quote, and it's a pretty damning one for their own method.
Jane: But they also found something interesting. The two best fixed estimators, the Beta-Binomial and the target-only logistic model, they were almost tied. The difference between them was zero point zero zero zero three on the Brier score.
Tom: And that's smaller than the difference you get just by changing the software version you use to fit the model. So the paper says you can't recommend one over the other. They're statistically indistinguishable.
Jane: So the improvement they're suggesting is not "use this specific estimator." It's "follow this process, do the gate, do the local labeling, and be humble about your ability to pick the best tool."
Tom: And that humility is the real takeaway. The paper is saying, don't trust a model's performance at another hospital, don't trust your ability to pick the perfect estimator, and definitely don't trust the uncertainty intervals they give you.
Jane: Because those intervals, the ninety-five percent ones, they only covered the truth eighty-seven percent of the time. And it was worse at the hardest hospital, MIMIC-CXR, where it dropped to seventy-eight percent.
Tom: So the improvements are really about process, not about a magic estimator. Next, we're going to look at the first page of the paper itself, and see how they set up this whole investigation.
First Page: Tom: So Jane, we've talked about the results, we've talked about the method. Let's go back to the very beginning. The first page of "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions" sets up the whole problem.
Jane: And the setup is really about a gap. The models are being sold as ready for clinical use, but the interfaces they provide don't give you any confidence score. You ask "is there pneumonia?" and it just says "yes" or "no." No probability, no "I'm eighty percent sure."
Tom: Right, and that's a problem for a hospital. Because if you get a "yes" from the model, you need to know how much to trust that "yes." Is it a strong signal, or is it the model just guessing?
Jane: The paper calls this the "estimand." And they're very careful to define it. They're measuring agreement with the institution's reference standard. Not clinical correctness. Not "is the patient actually sick." Just "does the model agree with what this hospital's records say?"
Tom: And that's a crucial distinction, because a hospital's labels can be wrong. The paper even did a small audit, one radiologist reviewing one hundred fifty studies, and the agreement with the dataset labels was only a kappa of zero point five zero two. That's moderate agreement at best.
Jane: So the labels themselves are noisy. And the paper is saying, we're not judging the model against truth. We're judging it against a noisy, imperfect reference standard. And that's the only thing a hospital can actually measure.
Tom: And the first page also introduces the three datasets. MIMIC-CXR, which is from an intensive care population, high acuity, high prevalence of disease. OpenI, which is a more routine outpatient setting. And PadChest, which is from a Spanish hospital, with labels in a different language.
Jane: And those differences matter. The prevalence of findings ranges from zero point one nine to zero point nine seven at MIMIC-CXR, but only zero point zero zero four to zero point zero nine seven at OpenI. That's a huge difference in what the model is seeing.
Tom: And that's the core of the external validation problem. If you train or test a model on a high-prevalence population, and then deploy it in a low-prevalence population, the error patterns are going to be completely different.
Jane: The first page also mentions the three models they used. MedGemma, which is a general medical model. CheXagent, which is a specialist for chest X-rays. And LLaVA-Med, which is a general vision-language model.
Tom: And they're all different sizes, different architectures, different training data. So they're testing whether the problems they find are specific to one model or general across all of them.
Jane: And spoiler alert, the problems were general. The heterogeneity, the inconsistency, the silence of some models at some sites, it wasn't just one bad apple. It was a pattern across all of them.
Tom: So the first page sets up a rigorous, multi-site, multi-model investigation. And it promises to be honest about the limitations. Next, we're going to wrap this up and talk about what it all means.
Conclusion: Tom: Alright, we've spent a lot of time with "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." Let's bring it home.
Jane: And the home message is a mixed one. The paper is incredibly thorough. Over three hundred forty-five thousand finding-level predictions, three models, three institutions, six findings, two elicitation protocols. That's a massive amount of evidence.
Tom: And the evidence says that these models are not consistent. Their error patterns change dramatically depending on the hospital, the finding, and even how you ask the question. One model went from "nothing is wrong" to "everything is wrong" just by changing the prompt.
Jane: And the paper's proposed solution, which is to have each hospital measure the model's agreement locally, is sound in principle. But the measurement tools they evaluated didn't always work. The adaptive selection failed, and the uncertainty intervals undercovered.
Tom: So what's the takeaway for a hospital? It's that you can't skip the work. You can't rely on a model's reported performance at another site. You have to label your own data, run your own tests, and be humble about what you can conclude.
Jane: And even then, the paper says, the two best estimators are too close to call. The difference between them is smaller than the software version sensitivity. So you can't even be sure you're using the best tool.
Tom: But the paper is also careful to say what it's not saying. It's not saying these models are clinically useless. It's saying that agreement with institutional labels is not the same as clinical correctness. And that's a distinction that gets lost a lot in the hype.
Jane: And that's the legacy of this paper, I think. It's a call for rigor. A call for hospitals to audit these models the way they'd audit any other medical device. And a call for researchers to be honest about the limits of their methods.
Tom: So we're saying goodbye to this paper, but the conversation is far from over. The next paper on our list is going to push these ideas even further, I hope.
Jane: And we'll be here to break it down. Thanks for listening, everyone. We'll see you on the next episode.
Tom: Take care, folks.
Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi
University College Dublin · The Third Affiliated Hospital of Southern Medical University · Dublin City University
cs.CV, cs.LG
Submitted: 2026-08-01
Comments: 10 pages, 3 figures, 3 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 58/100
The gist: The paper "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions" investigates whether agreement with an institution's reference standard
Terminology
Summary
The paper Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions
investigates whether agreement with an institution's reference standard for chest-radiograph findings from generative vision-language models (VLMs) transfers across sites, findings, prediction directions, and question formats. The authors evaluated three generative VLMs—MedGemma 4B, CheXagent 8B, and LLaVA-Med 7B—on three institutional chest-radiograph corpora (MIMIC-CXR, OpenI, and PadChest) across six findings (atelectasis, cardiomegaly, consolidation, edema, pleural effusion, pneumothorax) under two elicitation protocols (narrative and binary), comprising more than 345,000 finding-level predictions. The study estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels (25–200 labels), with estimation strategies stress-tested under repeated strict institution-held-out evaluation.
The paper makes three contributions. First, it characterizes how reference agreement varies with model, institution, finding, prediction direction, and elicitation protocol: after Benjamini–Hochberg adjustment over all 36 contrasts per metric, 23 false-positive-rate and 24 false-negative-rate contrasts remain significant, the institution effect persists with the finding held fixed, and the two primary VLMs, on identical images and labels, agree on about a third of the significant contrasts (Jaccard 0.35). Second, it provides a strict external validation of estimation strategies under a design excluding the receiving institution from development altogether: "adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta–Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Third, it presents an uncertainty audit treating interval coverage as an endpoint:
an interval aligned with what a held-out fold can test does not attain its nominal level, the shortfall concentrating at the hardest institution and in predictions asserting a finding."
Key findings include: "Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. The paper concludes:
Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness."
The paper also documents that LLaVA-Med contributed no false-positive contrast at all, having asserted no finding as present in any of its 60,070 evaluable judgments
under narrative elicitation, while under binary elicitation it asserted every finding as present. The authors note that the two leading policies change places between the interfaces
in a post-hoc matched five-candidate sensitivity analysis. The study's limitations include that the evidence comes from three public retrospective corpora rather than a live service, so institution-level inference rests on three clusters, too few for any interval here to be confirmatory,
and that what is estimated is agreement with those labels on the subset carrying a definite label; reader–dataset agreement was κ = 0.502, so 97.0% or 60.2% is agreement with a label of moderate quality, not with clinical truth.
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
Improvement: Build a post-hoc reliability estimator that works when the VLM exposes no confidence scores. Use the Beta–Binomial empirical-Bayes approach (Equation 2–3) as the default, with a target-only logistic model as a tied alternative.
What the improved system can do: Given a small budget of 25–200 labeled judgments from a receiving institution, it estimates per-finding, per-direction agreement with that institution's reference standard—without needing model internals, logits, or retraining. It achieves mean Brier 0.0853 and mean ECE 0.054 across 24 external-validation settings, outperforming pooled estimators by up to 0.084 in calibration error at the hardest site.
Improvement: Add a validation step that compares candidate estimators against a constant predictor under leave-one-institution-out cross-validation. If no estimator beats the constant, report the constant and stop.
Improvement: Detect whether the VLM is being asked via narrative (free-text) or binary (yes/no) elicitation, and switch estimators accordingly. The paper shows the two leading policies change places between protocols.
Improvement: Monitor the marginal assertion rate against reference prevalence. If a model's assertion rate is constant (e.g., always present or never present) across varying prevalence, flag it as degenerate.
Improvement: Instead of reporting credible intervals on latent probabilities, report the exact equal-tailed 95% posterior-predictive interval on the held-out count, using Beta-Binomial with the source-informed prior.
Improvement: When fitting logistic-family estimators, detect perfect separation in interaction indicators and report the coefficient sensitivity to solver implementation.
Improvement: Extend the categorical cell estimator with token log-probabilities from the affirmative/negative response tokens, enabling within-cell ranking.
Improvement: When sampling target labels, verify that no study appears in both fit and test sets, and report patient-level clustering if identifiers are available.
Summary of what the improved AI system can do: It reliably estimates reference agreement at a new institution with a small label budget, adapts to the elicitation protocol, detects degenerate models, avoids overconfident intervals, and flags solver sensitivity—all without requiring model internals or retraining.
Sources
- MedGemma Technical Report
- A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models