Performance of large language models in the optical diagnosis of colorectal polyps

summary

Video file (mp4)

The gist

Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis.

In short

The episode discusses a paper testing five large language models for diagnosing colorectal polyps from optical images. Models performed well at distinguishing neoplastic from non-neoplastic polyps, but struggled with identifying invasive carcinoma and did not meet clinical standards. The authors suggest human-in-the-loop workflows and need better data before clinical use.

Key concepts

Large Language Models (LLMs)
These are general-purpose chatbots used in this study to see if they can perform the specialized task of diagnosing colorectal polyps from pictures, testing whether a generalist AI can do a specialist's job.
F1 Score
This is a metric used to measure how well the models classified polyps. A high F1 score indicates strong performance in classifying them according to different systems like the Paris or NICE classification.
Human-in-the-loop workflows
The suggested improvement involves pairing AI with trainee endoscopists. This allows trainees to learn from the AI's reasoning, using the model's explanation as a training tool rather than just a diagnostic output.

Terminology used across episodes

This episode discusses

The paper

Performance of large language models in the optical diagnosis of colorectal polyps · Read on arXiv

Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V.G.L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles Ménard, Joshua Landy, Samir C. Grover

Scarborough Health Network Research Institute · University of Toronto · The Hospital for Sick Children · University of Calgary · Queen's University · Kingston Health Sciences Centre · Université de Montréal · Centre de Recherche du Centre hospitalier de l’Université de Montréal · University Hospital Würzburg · Mayo Clinic · Université de Sherbrooke · Centre Hospitalier Universitaire de Sherbrooke

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Performance of large language models in the optical diagnosis of colorectal polyps".

Jane: The paper was written by Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan et al. from Scarborough Health Network Research Institute and University of Toronto and The Hospital for Sick Children and University of Calgary and Queen's University and Kingston Health Sciences Centre and Université de Montréal and Centre de Recherche du Centre hospitalier de l’Université de Montréal and University Hospital Würzburg and Mayo Clinic and Université de Sherbrooke and Centre Hospitalier Universitaire de Sherbrooke.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Jane, we're looking at a paper that's got a real mouthful of a title today: "Performance of large language models in the optical diagnosis of colorectal polyps."

Jane: And honestly, Tom, that title is doing a lot of work. We're talking about whether the same kind of AI that writes emails and summarizes documents can actually look at a picture of a polyp in your bowel and tell your doctor whether it's dangerous.

Tom: Right, and that's a big deal because during a colonoscopy, doctors have to make a split-second call. Do we snip this thing out? Do we leave it alone? When does the patient need to come back?

Jane: Exactly. And the paper is testing five different large language models to see if they can make that call as well as an expert would. The authors are from a bunch of Canadian hospitals, plus some folks in Germany and the US.

Tom: And the key thing here is that these aren't specialized medical tools. These are the general-purpose chatbots you and I could go use right now. The researchers are asking: can a generalist AI do a specialist's job?

Jane: That's the exciting part, Tom. If a general chatbot can do this, it changes who has access to expert-level polyp diagnosis. You don't need a million-dollar system. You might just need an API key.

Tom: But, and I suspect this is where it gets interesting, the paper is probably going to show us that it's not quite that simple. Let's get into what they actually found.

Jane: Let's do it.

Summary: Tom: So Jane, we've got the title. Let's talk about what this paper actually did, because the summary is pretty fascinating. They took one hundred thirty-two cases of colorectal polyps, each with both white light and narrow-band imaging pictures.

Jane: And they showed those images to five different models. We've got Claude Opus four Google's Gemini two point five Pro, and three versions of OpenAI's GPT models — the o3, the 4o, and the brand new GPT-five.

Tom: Now, the big question was: can these models classify polyps the way a trained endoscopist would? They tested three classification systems. The Paris system, which describes the shape of the polyp. The NICE system, which looks at color and surface patterns. And then predicted histology — basically, what would the lab find if they biopsied it?

Jane: And here's the headline result, Tom. For telling neoplastic polyps apart from non-neoplastic ones — that's the difference between something that could become cancer and something harmless — all five models scored F1 scores above zero point nine. That's really strong.

Tom: But then you dig a little deeper and the story changes. When they tried to identify invasive carcinoma, the really dangerous stuff, the models struggled. Gemini two point five Pro got an F1 score of zero point five six zero. Some of the others scored zero.

Jane: Zero is rough. That means they never got it right.

Tom: Right. And the accuracy for the Paris classification was all over the place. Claude Opus four and GPT-five both hit forty-one point seven percent correct, which sounds low, but it was still statistically better than the other models.

Jane: So the takeaway isn't that these models are useless. It's that they're good at some things and bad at others. And that distinction matters a lot when you're talking about clinical use.

Tom: Exactly. And it's why the authors are careful to say these results don't meet the standards set by the European Society of Gastrointestinal Endoscopy. We'll get into what that means next.

Improvements: Jane: Tom, let's talk about what the paper suggests we should actually do with these findings, because the authors aren't just saying "these models are good" or "these models are bad." They're proposing a path forward.

Tom: And that path involves something called human-in-the-loop workflows. The idea is that you don't let the AI make the call alone. You pair it with a trainee endoscopist who's still learning.

Jane: Right, and that's clever because the models are actually pretty good at explaining their reasoning. Unlike a traditional computer-aided diagnosis system that just puts a heat map on the screen, these large language models can say "I think this is a Type two NICE lesion because the surface pattern looks regular and the color is uniform."

Tom: So a trainee can learn from that reasoning. They can compare their own thinking to the model's thinking. That's a training tool, not just a diagnostic tool.

Jane: And the paper also digs into something practical: how you write the prompt matters. They tested a long, structured prompt with definitions and a decision algorithm against a short, simple prompt. And the results were mixed.

Tom: Mixed is the right word. Claude Opus four and Gemini two point five Pro did better with the long prompt for the NICE classification. But GPT-o3, GPT-4o, and GPT-five actually did better with the short prompt.

Jane: Which is a reminder that these models are not all the same. They're built differently, they reason differently, and they respond differently to instructions. You can't just assume one prompt works for all of them.

Tom: And the authors also flag that the dataset itself has limitations. There aren't enough hyperplastic polyps in there, and the images are single views. Real colonoscopies give you multiple angles.

Jane: So the improvement they're really calling for is better data and more careful study design. They want prospective multicenter trials before anyone uses these in a clinic.

Tom: And that brings us to the actual first page of the paper, where they lay out the stakes.

First Page: Tom: So Jane, let's go back to the very first page of "Performance of large language models in the optical diagnosis of colorectal polyps" and look at how they frame the problem.

Jane: And the framing is really important, Tom. The first page tells us that colonoscopy prevents colorectal cancer by finding and removing precancerous polyps. That's the whole point of the procedure.

Tom: But here's the catch they highlight: optical diagnosis — looking at a polyp and deciding what it is — is hard. Even experts struggle with it. The paper cites studies showing that endoscopists are inadequately trained in recognizing T1 colorectal carcinomas, which are the early invasive ones.

Jane: And that's where AI comes in. They mention computer-aided detection and computer-aided diagnosis systems that already exist. But those systems have a problem: they don't generalize well across institutions and imaging setups.

Tom: So the argument is that large language models might be different. They're flexible, they can reason, and they can explain themselves. The question is whether that flexibility translates into accuracy.

Jane: And the first page also tells us who's behind this. It's a big collaborative team led by Samir Grover at the Scarborough Health Network in Toronto. They're using a dataset called PRIME, which was originally built to train endoscopists, not to test AI.

Tom: That's a key detail. The dataset wasn't designed for this study. It was designed for human education. So using it to test AI is a bit of a stretch, and the authors acknowledge that.

Jane: Right. And they also mention the guidelines they're measuring against. The ASGE and ESGE have specific thresholds for what makes optical diagnosis good enough to use in practice. The ESGE wants sensitivity of at least ninety percent and specificity of at least eighty percent.

Tom: And the models didn't hit those numbers. That's the honest bottom line from the first page — the promise is there, but the performance isn't there yet.

Jane: But it's close enough that the authors think it's worth pursuing. And that's a big deal for where this field is heading.

Conclusion: Tom: Alright Jane, let's wrap this up. We've been talking about "Performance of large language models in the optical diagnosis of colorectal polyps" and I think the picture is pretty clear now.

Jane: It is, Tom. The models are genuinely impressive at telling neoplastic from non-neoplastic polyps — that's the cancer-versus-not-cancer question. But they fall short on the harder tasks, like spotting invasive carcinoma or grading dysplasia.

Tom: And the authors are honest about that. They say the performance doesn't meet the ESGE standards for clinical use. But they also point out that these models beat trainee-level performance in some areas, and they can explain their reasoning in plain language.

Jane: Which is why they're pushing for human-in-the-loop systems. Pair the AI with a trainee, let the trainee learn from the AI's reasoning, and you might get better outcomes than either could achieve alone.

Tom: The other big takeaway is the prompting finding. Long structured prompts helped some models and hurt others. That's a reminder that we're still figuring out how to talk to these systems.

Jane: And the dataset limitations are real. Single-view images, not enough hyperplastic polyps, no size or location data. The authors want prospective multicenter trials before anyone uses this in a clinic.

Tom: So the verdict is: promising, but not ready. And that's actually a healthy place for the research to be.

Jane: Agreed. It's a solid study with clear limitations and a clear path forward. We'll be watching to see if the next generation of models does better.

Tom: And that's a wrap on this one. Thanks for listening, and we'll see you on the next paper.

Jane: See you soon, everyone.

More episodes

← Home