Performance of large language models in the optical diagnosis of colorectal polyps
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Performance of large language models in the optical diagnosis of colorectal polyps".
Jane: The paper was written by Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan et al. from Scarborough Health Network Research Institute and University of Toronto and The Hospital for Sick Children and University of Calgary and Queen's University and Kingston Health Sciences Centre and Université de Montréal and Centre de Recherche du Centre hospitalier de l’Université de Montréal and University Hospital Würzburg and Mayo Clinic and Université de Sherbrooke and Centre Hospitalier Universitaire de Sherbrooke.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Jane, we're looking at a paper that's got a real mouthful of a title today: "Performance of large language models in the optical diagnosis of colorectal polyps."
Jane: And honestly, Tom, that title is doing a lot of work. We're talking about whether the same kind of AI that writes emails and summarizes documents can actually look at a picture of a polyp in your bowel and tell your doctor whether it's dangerous.
Tom: Right, and that's a big deal because during a colonoscopy, doctors have to make a split-second call. Do we snip this thing out? Do we leave it alone? When does the patient need to come back?
Jane: Exactly. And the paper is testing five different large language models to see if they can make that call as well as an expert would. The authors are from a bunch of Canadian hospitals, plus some folks in Germany and the US.
Tom: And the key thing here is that these aren't specialized medical tools. These are the general-purpose chatbots you and I could go use right now. The researchers are asking: can a generalist AI do a specialist's job?
Jane: That's the exciting part, Tom. If a general chatbot can do this, it changes who has access to expert-level polyp diagnosis. You don't need a million-dollar system. You might just need an API key.
Tom: But, and I suspect this is where it gets interesting, the paper is probably going to show us that it's not quite that simple. Let's get into what they actually found.
Jane: Let's do it.
Summary: Tom: So Jane, we've got the title. Let's talk about what this paper actually did, because the summary is pretty fascinating. They took one hundred thirty-two cases of colorectal polyps, each with both white light and narrow-band imaging pictures.
Jane: And they showed those images to five different models. We've got Claude Opus four Google's Gemini two point five Pro, and three versions of OpenAI's GPT models — the o3, the 4o, and the brand new GPT-five.
Tom: Now, the big question was: can these models classify polyps the way a trained endoscopist would? They tested three classification systems. The Paris system, which describes the shape of the polyp. The NICE system, which looks at color and surface patterns. And then predicted histology — basically, what would the lab find if they biopsied it?
Jane: And here's the headline result, Tom. For telling neoplastic polyps apart from non-neoplastic ones — that's the difference between something that could become cancer and something harmless — all five models scored F1 scores above zero point nine. That's really strong.
Tom: But then you dig a little deeper and the story changes. When they tried to identify invasive carcinoma, the really dangerous stuff, the models struggled. Gemini two point five Pro got an F1 score of zero point five six zero. Some of the others scored zero.
Jane: Zero is rough. That means they never got it right.
Tom: Right. And the accuracy for the Paris classification was all over the place. Claude Opus four and GPT-five both hit forty-one point seven percent correct, which sounds low, but it was still statistically better than the other models.
Jane: So the takeaway isn't that these models are useless. It's that they're good at some things and bad at others. And that distinction matters a lot when you're talking about clinical use.
Tom: Exactly. And it's why the authors are careful to say these results don't meet the standards set by the European Society of Gastrointestinal Endoscopy. We'll get into what that means next.
Improvements: Jane: Tom, let's talk about what the paper suggests we should actually do with these findings, because the authors aren't just saying "these models are good" or "these models are bad." They're proposing a path forward.
Tom: And that path involves something called human-in-the-loop workflows. The idea is that you don't let the AI make the call alone. You pair it with a trainee endoscopist who's still learning.
Jane: Right, and that's clever because the models are actually pretty good at explaining their reasoning. Unlike a traditional computer-aided diagnosis system that just puts a heat map on the screen, these large language models can say "I think this is a Type two NICE lesion because the surface pattern looks regular and the color is uniform."
Tom: So a trainee can learn from that reasoning. They can compare their own thinking to the model's thinking. That's a training tool, not just a diagnostic tool.
Jane: And the paper also digs into something practical: how you write the prompt matters. They tested a long, structured prompt with definitions and a decision algorithm against a short, simple prompt. And the results were mixed.
Tom: Mixed is the right word. Claude Opus four and Gemini two point five Pro did better with the long prompt for the NICE classification. But GPT-o3, GPT-4o, and GPT-five actually did better with the short prompt.
Jane: Which is a reminder that these models are not all the same. They're built differently, they reason differently, and they respond differently to instructions. You can't just assume one prompt works for all of them.
Tom: And the authors also flag that the dataset itself has limitations. There aren't enough hyperplastic polyps in there, and the images are single views. Real colonoscopies give you multiple angles.
Jane: So the improvement they're really calling for is better data and more careful study design. They want prospective multicenter trials before anyone uses these in a clinic.
Tom: And that brings us to the actual first page of the paper, where they lay out the stakes.
First Page: Tom: So Jane, let's go back to the very first page of "Performance of large language models in the optical diagnosis of colorectal polyps" and look at how they frame the problem.
Jane: And the framing is really important, Tom. The first page tells us that colonoscopy prevents colorectal cancer by finding and removing precancerous polyps. That's the whole point of the procedure.
Tom: But here's the catch they highlight: optical diagnosis — looking at a polyp and deciding what it is — is hard. Even experts struggle with it. The paper cites studies showing that endoscopists are inadequately trained in recognizing T1 colorectal carcinomas, which are the early invasive ones.
Jane: And that's where AI comes in. They mention computer-aided detection and computer-aided diagnosis systems that already exist. But those systems have a problem: they don't generalize well across institutions and imaging setups.
Tom: So the argument is that large language models might be different. They're flexible, they can reason, and they can explain themselves. The question is whether that flexibility translates into accuracy.
Jane: And the first page also tells us who's behind this. It's a big collaborative team led by Samir Grover at the Scarborough Health Network in Toronto. They're using a dataset called PRIME, which was originally built to train endoscopists, not to test AI.
Tom: That's a key detail. The dataset wasn't designed for this study. It was designed for human education. So using it to test AI is a bit of a stretch, and the authors acknowledge that.
Jane: Right. And they also mention the guidelines they're measuring against. The ASGE and ESGE have specific thresholds for what makes optical diagnosis good enough to use in practice. The ESGE wants sensitivity of at least ninety percent and specificity of at least eighty percent.
Tom: And the models didn't hit those numbers. That's the honest bottom line from the first page — the promise is there, but the performance isn't there yet.
Jane: But it's close enough that the authors think it's worth pursuing. And that's a big deal for where this field is heading.
Conclusion: Tom: Alright Jane, let's wrap this up. We've been talking about "Performance of large language models in the optical diagnosis of colorectal polyps" and I think the picture is pretty clear now.
Jane: It is, Tom. The models are genuinely impressive at telling neoplastic from non-neoplastic polyps — that's the cancer-versus-not-cancer question. But they fall short on the harder tasks, like spotting invasive carcinoma or grading dysplasia.
Tom: And the authors are honest about that. They say the performance doesn't meet the ESGE standards for clinical use. But they also point out that these models beat trainee-level performance in some areas, and they can explain their reasoning in plain language.
Jane: Which is why they're pushing for human-in-the-loop systems. Pair the AI with a trainee, let the trainee learn from the AI's reasoning, and you might get better outcomes than either could achieve alone.
Tom: The other big takeaway is the prompting finding. Long structured prompts helped some models and hurt others. That's a reminder that we're still figuring out how to talk to these systems.
Jane: And the dataset limitations are real. Single-view images, not enough hyperplastic polyps, no size or location data. The authors want prospective multicenter trials before anyone uses this in a clinic.
Tom: So the verdict is: promising, but not ready. And that's actually a healthy place for the research to be.
Jane: Agreed. It's a solid study with clear limitations and a clear path forward. We'll be watching to see if the next generation of models does better.
Tom: And that's a wrap on this one. Thanks for listening, and we'll see you on the next paper.
Jane: See you soon, everyone.
Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V.G.L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles Ménard, Joshua Landy, Samir C. Grover
Scarborough Health Network Research Institute · University of Toronto · The Hospital for Sick Children · University of Calgary · Queen's University · Kingston Health Sciences Centre · Université de Montréal · Centre de Recherche du Centre hospitalier de l’Université de Montréal · University Hospital Würzburg · Mayo Clinic · Université de Sherbrooke · Centre Hospitalier Universitaire de Sherbrooke
cs.CV, cs.AI
Submitted: 2026-07-31
Updated: 2026-08-11
Comments: 22 pages, 1 figure, 5 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 57/100
The gist: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis.
Key concepts
- Large Language Models (LLMs)
- These are general-purpose chatbots used in this study to see if they can perform the specialized task of diagnosing colorectal polyps from pictures, testing whether a generalist AI can do a specialist's job.
- F1 Score
- This is a metric used to measure how well the models classified polyps. A high F1 score indicates strong performance in classifying them according to different systems like the Paris or NICE classification.
- Human-in-the-loop workflows
- The suggested improvement involves pairing AI with trainee endoscopists. This allows trainees to learn from the AI's reasoning, using the model's explanation as a training tool rather than just a diagnostic output.
Terminology
Summary
Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. The authors aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology.
This was a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. The authors evaluated five multimodal MLLMs: Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology classifications, they calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM.
The study followed the STARD 2015 guidelines for diagnostic accuracy studies. Ethics approval was obtained from the Scarborough Health Network Research Ethics Board (REB# GAS-25-012). Each case was evaluated ten times per model with an identical prompt and fixed parameters, with the most frequent diagnostic classification across replicates used as the primary predicted classification. Two prompt modalities were designed: a long, structured prompt including role and objective specification, warnings, a definitions primer, a clinical decision-making algorithm, and a strict JSON output format; and a short prompt with a brief role and output schema.
The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification.
For Paris classification, the percentage of correct responses ranged from 24.2% to 41.7%, with Cochran's Q of 32.9 (p < 0.001), indicating a statistically significant difference between models. Claude Opus 4 and GPT-5 demonstrated statistically significantly higher accuracy (41.7%) than Gemini 2.5 Pro, GPT-o3, and GPT-4o, with scores of 30.3%, 33.3%, and 24.2%, respectively.
For NICE classification, the percentage of correct responses ranged from 78.0% to 84.1%, with Cochran's Q of 4.48 (p=0.345), indicating no statistically significant difference between the five MLLMs. F1 scores for invasive vs. non-invasive classifications ranged from 0.00 to 0.560, while accuracy ranged from 91.7% to 94.7%. F1 scores for neoplastic vs. non-neoplastic classifications ranged from 0.910 to 0.937, with accuracy ranging from 84.1% to 88.6%.
For predicted histology, the percentage of correct responses ranged from 55.3% to 62.9%, with Cochran's Q of 4.19 (p=0.381), indicating no statistically significant difference between the MLLMs. F1 scores for invasive vs. non-invasive classifications ranged from 0.00 to 0.500, with accuracy ranging from 92.4% to 94.7%. F1 scores for neoplastic vs. non-neoplastic classifications ranged from 0.965 to 0.981, with accuracy ranging from 93.2% to 96.2%. F1 scores for low-grade vs. high-grade classifications ranged from 0.00 to 0.492, with accuracy ranging from 67.0% to 75.5%.
For the short vs. long prompt modality analysis, with NICE, Claude Opus 4 and Gemini 2.5 Pro achieved higher accuracy with the long prompt (84.8% and 78.8% for the long prompt vs 81.8% and 73.5% for the short prompt respectively, both McNemar p<0.001), whereas GPT-o3/4o/5 performed better with the short prompt (all p<0.001).
Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment. The authors concluded that "MLLMs show substantial promise in the optical diagnosis of colorectal polyps. These results justify prospective, multi center trials and careful human-in-the-loop design to consider their potential use for clinical practice and endoscopy training."
The PRIME dataset is retrospective, which constrains generalizability. The dataset was not curated for a diagnostic accuracy study, and certain polyp types, such as hyperplastic and Paris 0-III, lack sufficient representation. Additionally, because the dataset was curated for human evaluation, some images lack the quality ideal for MLLM diagnoses. Images offer a single view of the polyp and lack size and location metadata. MLLM parameters, such as temperature and top p, may vary across providers.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system for optical diagnosis of colorectal polyps:
-
Improvement: Instead of relying on a single MLLM, build an ensemble that dynamically selects the best model per classification task based on the paper's findings.
-
What it can do:
-
Use Gemini 2.5 Pro for invasive vs. non-invasive classification (highest F1: 0.560) and low- vs. high-grade adenoma (F1: 0.492).
-
Use Claude Opus 4 for neoplastic vs. non-neoplastic differentiation (highest F1: 0.937 for NICE, 0.968 for histology).
-
Use GPT-5 or Claude Opus 4 for Paris classification (both at 41.7% accuracy, significantly higher than others).
-
Route each case to the optimal model, improving overall diagnostic accuracy beyond any single model.
-
Improvement: Implement an adaptive prompt selection mechanism based on the paper's finding that prompt length affects performance differently across models.
-
What it can do:
-
Automatically use long structured prompts for Claude Opus 4 and Gemini 2.5 Pro (improved NICE accuracy from 81.8% to 84.8% and 73.5% to 78.8% respectively).
-
Use short prompts for GPT-o3, GPT-4o, and GPT-5 (all performed better with short prompts for NICE).
-
For Paris classification, use long prompts for GPT-5 (improved from 31.8% to 41.7%).
-
This dynamic prompting could improve accuracy by 3-10 percentage points depending on the model-task combination.
-
Improvement: Add a confidence threshold system that flags low-confidence predictions for human review, addressing the paper's finding that models struggle with rare classes.
-
What it can do:
-
Detect when the model is likely biased toward common classifications (e.g., NICE Type 2) and flag cases where invasive carcinoma is suspected but model confidence is low.
-
Automatically escalate uncertain cases to a human endoscopist, preventing missed diagnoses of high-risk lesions.
-
This addresses the critical gap where models showed F1 scores of 0.00 for invasive carcinoma detection in some models.
-
Improvement: Develop a system that integrates multiple image perspectives, addressing the paper's finding that the PRIME dataset's single-view limitation may have reduced performance.
-
What it can do:
-
Accept multiple images of the same polyp from different angles (as in the Massimi et al. study) and fuse the information before classification.
-
This could improve Paris classification accuracy, as the paper noted that multi-frame analysis in other datasets yielded better results.
-
Improvement: Implement oversampling or synthetic data augmentation for underrepresented classes (hyperplastic polyps, Paris 0-III, invasive carcinomas).
-
What it can do:
-
The paper noted only 5 hyperplastic polyps in 132 cases, contributing to poor specificity (as low as 0% for some models).
-
By training or fine-tuning on balanced datasets or using class-weighted loss functions, the system can improve specificity without sacrificing sensitivity.
-
Improvement: Create an interactive training tool that pairs the MLLM with trainee endoscopists, leveraging the paper's suggestion for novel workflows.
-
What it can do:
-
Allow trainees to input supplementary information (colonic segment, polyp size) that the model can use to refine its diagnosis.
-
Provide explainable reasoning (unlike traditional CADx heat maps) to help trainees understand why a polyp is classified a certain way.
-
Track trainee performance against the MLLM and provide targeted feedback on specific classification errors.
-
Improvement: Fine-tune the best-performing model (Gemini 2.5 Pro) specifically on NICE classification tasks with a focus on Type 3 (invasive) detection.
-
What it can do:
-
The paper showed Gemini 2.5 Pro achieved 70% sensitivity for invasive carcinoma but only 46.7% PPV.
-
Fine-tuning on a larger dataset of invasive polyps could improve precision while maintaining recall, bringing performance closer to ESGE guidelines (≥90% sensitivity, ≥80% specificity).
-
Improvement: Run multiple replicates (as the paper did with 10 per case) and use inter-replicate agreement as a confidence metric.
-
What it can do:
-
Flag cases where the model gives inconsistent answers across replicates (indicating uncertainty).
-
For high-agreement cases, the system can provide diagnoses with higher confidence.
-
This addresses the paper's finding that McNemar's tests showed asymmetric disagreements even when overall accuracy was similar.
-
Improvement: Build a system that outputs not just classifications but also recommended actions aligned with ESGE/ASGE guidelines.
-
What it can do:
-
Automatically suggest
resect-and-discard
ordiagnose-and-leave
strategies when model confidence meets guideline thresholds. -
Flag cases where the model's performance falls below ESGE standards (sensitivity <90%, specificity <80%) and recommend pathology evaluation instead.
-
This bridges the gap between raw classification and clinical decision-making.
-
Improvement: Implement a feedback loop where expert-verified diagnoses are used to continuously improve the system.
-
What it can do:
-
Collect expert consensus diagnoses (as in the PRIME dataset's three-expert committee) and use them to periodically fine-tune the models.
-
Track performance drift over time and alert when model accuracy drops below acceptable thresholds.
-
This addresses the paper's call for prospective multicenter trials and ensures the system remains clinically relevant.
Sources
- Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems
- Effects of Prompt Length on Domain-specific Tasks for Large Language Models
- From System 1 to System 2: A Survey of Reasoning Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models