Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

summary

Video file (mp4)

The gist

The paper "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions" investigates whether agreement with an institution's reference standard

This episode discusses

The paper

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions · Read on arXiv

Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi

University College Dublin · The Third Affiliated Hospital of Southern Medical University · Dublin City University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions".

Jane: The paper was written by Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng et al. from University College Dublin and The Third Affiliated Hospital of Southern Medical University and Dublin City University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everybody. Today we're looking at a paper that's got a real mouthful of a title: "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions."

Jane: And Tom, that title is actually doing a lot of work. Let me break it down. "Vision-language models" are those AI systems that can look at an image and then answer questions about it in text. So here, they're looking at chest X-rays.

Tom: Right, and the models are saying things like "this patient has fluid in their lungs" or "no sign of collapse." But the key word in that title is "auditing." These models are being checked, like a financial audit, to see if their answers actually hold up.

Jane: And the "across institutions" part is the twist. A model might be great at one hospital, but when you move it to a different hospital, with different patients, different equipment, different radiologists writing the reports, does it still get things right?

Tom: So it's not just "is the AI smart?" It's "is the AI trustworthy when I bring it into my building?" And the paper found some pretty wild stuff. I mean, one model literally stopped saying anything was wrong at one hospital. It just said "no findings" for everything.

Jane: That's the kind of thing that would terrify a radiologist. You think you have a tool that's helping you, and it's actually just nodding along. But the paper is careful to say they're measuring agreement with the hospital's own labels, not whether the patient is actually sick.

Tom: Right, that's a crucial distinction. They're not saying "this model is clinically wrong." They're saying "this model disagrees with your hospital's established readings, and the disagreement pattern changes depending on where you are."

Jane: And that's why the title matters. It's not "are these models good?" It's "can you measure how much to trust them at your specific site?" That's a much more practical question for a hospital administrator.

Tom: Exactly. And the answer they found is, honestly, a bit uncomfortable. You can't just borrow trust from another hospital. You have to measure it yourself, with your own data.

Jane: So stick around, because we're going to get into how they actually tried to do that measurement, and whether their clever estimation tricks actually worked. Because spoiler alert, some of them didn't.

Tom: That's the hook, folks. The models are powerful, but trust doesn't transfer. Next up, we're going to dig into what the paper actually did and what they found.

Summary: Tom: So Jane, we've set the stage. The paper is "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." Now let's talk about what they actually did. This is a massive study.

Jane: Massive is the word. They took three different vision-language models, ran them on three different chest X-ray datasets, and looked at six different medical findings. That's like three different brands of X-ray glasses, tested on three different hospitals' patient populations.

Tom: And they didn't just ask the models once. They asked them two different ways. One way was a narrative prompt, like "describe what you see." The other was a binary prompt, like "is there edema? Yes or no?" And the way you ask matters. A lot.

Jane: That's one of the most striking findings, actually. The same model, on the same image, would give a different answer depending on how the question was phrased. We're talking about a kappa score of zero point two, which means the two protocols barely agreed with each other.

Tom: And it gets worse. One of the models, LLaVA-Med, under the narrative prompt, never said a single finding was present. Not once, in over sixty thousand judgments. But under the binary prompt, it said every single finding was present. It went from "nothing is wrong" to "everything is wrong" just by changing the question format.

Jane: That's a complete breakdown. And it tells you that these models aren't really "reading" the X-ray the way a human does. They're responding to the prompt structure in ways that are hard to predict.

Tom: But here's the thing, Jane. The paper isn't just documenting these failures. They're trying to solve a practical problem. If I'm a hospital, and I want to use one of these models, how do I know if I can trust it at my site?

Jane: And their answer is, you need to test it yourself. You need to take a small sample of your own X-rays, get your own radiologists to label them, and see how the model does against those labels. They call this a "target label budget."

Tom: So you spend maybe a hundred or two hundred X-rays, you label them yourself, and then you try to estimate the model's agreement rate at your hospital. And the paper compares ten different ways of doing that estimation.

Jane: And that's where the results get really interesting, because the clever methods didn't always win. We'll get into that in the next segment, but the short version is, sometimes the simple approach is the best approach.

Tom: And sometimes, no approach is good enough. The uncertainty intervals they built, the ones that are supposed to tell you "we're ninety-five percent sure the agreement is between here and here," they only covered the truth eighty-seven percent of the time.

Jane: So even when you do the work, even when you label your own data, the estimates can still be off. That's a sobering message for anyone who wants to deploy these models.

Tom: So the summary is: the models are inconsistent, the measurement is hard, and the trust has to be earned locally. Next, we're going to talk about the improvements the paper suggests, and whether they actually hold up.

Improvements: Tom: Welcome back. We're still on "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." And Jane, I want to dig into the "improvements" part of the paper. Because they don't just complain, they actually propose a method.

Jane: Right, they lay out a whole procedure. It's in the paper as Algorithm one. It's basically a step-by-step guide for a hospital to follow when they want to adopt one of these models.

Tom: And the first step is the most important one, I think. They call it a "fit-for-purpose gate." Before you even bother with fancy estimators, you check whether any estimator beats just using a single average number for everything.

Jane: And that's a really practical move. Because if the model is so bad that the best you can do is a constant "it's right about ninety percent of the time," then you don't need a complex statistical model. You just need that number.

Tom: And the paper found that this gate actually matters. In four out of two hundred and forty decisions, the gate correctly said "stop, just use the constant." That's not a huge number, but it shows the gate is doing something.

Jane: But here's where the improvements get messy. They designed a smart selection procedure. It looks at the available data, it ranks the different estimators, and it picks the best one for your site. And that sounds great in theory.

Tom: But in practice, it failed. The adaptive selection procedure had a Brier score of zero point one zero eight three. But if you just always used one simple estimator, the Beta-Binomial one, you got zero point zero eight five three. The adaptive procedure was worse than just picking a default and sticking with it.

Jane: And that's a really important negative result. It means that trying to be clever about which estimator to use, at least with the data available, actually hurt. The selection process itself introduced error.

Tom: Right, and the paper is honest about this. They say "adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives." That's a direct quote, and it's a pretty damning one for their own method.

Jane: But they also found something interesting. The two best fixed estimators, the Beta-Binomial and the target-only logistic model, they were almost tied. The difference between them was zero point zero zero zero three on the Brier score.

Tom: And that's smaller than the difference you get just by changing the software version you use to fit the model. So the paper says you can't recommend one over the other. They're statistically indistinguishable.

Jane: So the improvement they're suggesting is not "use this specific estimator." It's "follow this process, do the gate, do the local labeling, and be humble about your ability to pick the best tool."

Tom: And that humility is the real takeaway. The paper is saying, don't trust a model's performance at another hospital, don't trust your ability to pick the perfect estimator, and definitely don't trust the uncertainty intervals they give you.

Jane: Because those intervals, the ninety-five percent ones, they only covered the truth eighty-seven percent of the time. And it was worse at the hardest hospital, MIMIC-CXR, where it dropped to seventy-eight percent.

Tom: So the improvements are really about process, not about a magic estimator. Next, we're going to look at the first page of the paper itself, and see how they set up this whole investigation.

First Page: Tom: So Jane, we've talked about the results, we've talked about the method. Let's go back to the very beginning. The first page of "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions" sets up the whole problem.

Jane: And the setup is really about a gap. The models are being sold as ready for clinical use, but the interfaces they provide don't give you any confidence score. You ask "is there pneumonia?" and it just says "yes" or "no." No probability, no "I'm eighty percent sure."

Tom: Right, and that's a problem for a hospital. Because if you get a "yes" from the model, you need to know how much to trust that "yes." Is it a strong signal, or is it the model just guessing?

Jane: The paper calls this the "estimand." And they're very careful to define it. They're measuring agreement with the institution's reference standard. Not clinical correctness. Not "is the patient actually sick." Just "does the model agree with what this hospital's records say?"

Tom: And that's a crucial distinction, because a hospital's labels can be wrong. The paper even did a small audit, one radiologist reviewing one hundred fifty studies, and the agreement with the dataset labels was only a kappa of zero point five zero two. That's moderate agreement at best.

Jane: So the labels themselves are noisy. And the paper is saying, we're not judging the model against truth. We're judging it against a noisy, imperfect reference standard. And that's the only thing a hospital can actually measure.

Tom: And the first page also introduces the three datasets. MIMIC-CXR, which is from an intensive care population, high acuity, high prevalence of disease. OpenI, which is a more routine outpatient setting. And PadChest, which is from a Spanish hospital, with labels in a different language.

Jane: And those differences matter. The prevalence of findings ranges from zero point one nine to zero point nine seven at MIMIC-CXR, but only zero point zero zero four to zero point zero nine seven at OpenI. That's a huge difference in what the model is seeing.

Tom: And that's the core of the external validation problem. If you train or test a model on a high-prevalence population, and then deploy it in a low-prevalence population, the error patterns are going to be completely different.

Jane: The first page also mentions the three models they used. MedGemma, which is a general medical model. CheXagent, which is a specialist for chest X-rays. And LLaVA-Med, which is a general vision-language model.

Tom: And they're all different sizes, different architectures, different training data. So they're testing whether the problems they find are specific to one model or general across all of them.

Jane: And spoiler alert, the problems were general. The heterogeneity, the inconsistency, the silence of some models at some sites, it wasn't just one bad apple. It was a pattern across all of them.

Tom: So the first page sets up a rigorous, multi-site, multi-model investigation. And it promises to be honest about the limitations. Next, we're going to wrap this up and talk about what it all means.

Conclusion: Tom: Alright, we've spent a lot of time with "Auditing Medical Vision–Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions." Let's bring it home.

Jane: And the home message is a mixed one. The paper is incredibly thorough. Over three hundred forty-five thousand finding-level predictions, three models, three institutions, six findings, two elicitation protocols. That's a massive amount of evidence.

Tom: And the evidence says that these models are not consistent. Their error patterns change dramatically depending on the hospital, the finding, and even how you ask the question. One model went from "nothing is wrong" to "everything is wrong" just by changing the prompt.

Jane: And the paper's proposed solution, which is to have each hospital measure the model's agreement locally, is sound in principle. But the measurement tools they evaluated didn't always work. The adaptive selection failed, and the uncertainty intervals undercovered.

Tom: So what's the takeaway for a hospital? It's that you can't skip the work. You can't rely on a model's reported performance at another site. You have to label your own data, run your own tests, and be humble about what you can conclude.

Jane: And even then, the paper says, the two best estimators are too close to call. The difference between them is smaller than the software version sensitivity. So you can't even be sure you're using the best tool.

Tom: But the paper is also careful to say what it's not saying. It's not saying these models are clinically useless. It's saying that agreement with institutional labels is not the same as clinical correctness. And that's a distinction that gets lost a lot in the hype.

Jane: And that's the legacy of this paper, I think. It's a call for rigor. A call for hospitals to audit these models the way they'd audit any other medical device. And a call for researchers to be honest about the limits of their methods.

Tom: So we're saying goodbye to this paper, but the conversation is far from over. The next paper on our list is going to push these ideas even further, I hope.

Jane: And we'll be here to break it down. Thanks for listening, everyone. We'll see you on the next episode.

Tom: Take care, folks.

More episodes

← Home