Large language models improve physician accuracy but lead to false reliance
summary
The gist
The paper presents CORA (Citation-Oriented Retrieval Assistant), an agentic retrieval-augmented generation (RAG) system for dermatological decision support, and evaluates both its standalone
In short
The episode discusses a paper titled "Large language models improve physician accuracy but lead to false reliance." The hosts analyze how retrieval-augmented AI improves doctor accuracy but creates a risk where doctors over-rely on citations, even when the AI is wrong. The conclusion suggests systems must be designed to be more honest about uncertainty and flag weak evidence.
Key concepts
- False Reliance
- This occurs when doctors trust an AI's answer too much because the AI provides a citation, even if that citation does not actually support the answer. This leads doctors to follow incorrect advice.
- Grounding Miscalibration
- This is a problem where the signal provided by citations is correlated with correctness but is not strong enough to justify high levels of trust. Doctors treat supportive-looking citations as near-proof, even when they are only hints.
- Reliance-Stratified Reporting
- This proposed method for evaluating AI systems involves reporting performance separately based on whether the doctor and AI agree or disagree, especially focusing on cases where the AI is wrong to identify dangerous failures.
Terminology used across episodes
This episode discusses
The paper
Large language models improve physician accuracy but lead to false reliance · Read on arXiv
Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Titus J. Brinker
German Cancer Research Center · University Heidelberg · University Medical Center Mannheim · Ruprecht-Karl University of Heidelberg · Medical University of Vienna · Pontificia Universidad Católica de Chile · University Medical Center Rostock · Heidelberg University Hospital
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Large language models improve physician accuracy but lead to false reliance".
Jane: The paper was written by Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl et al. from German Cancer Research Center and University Heidelberg and University Medical Center Mannheim and Ruprecht-Karl University of Heidelberg and Medical University of Vienna and Pontificia Universidad Católica de Chile and University Medical Center Rostock and Heidelberg University Hospital.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's going to make you think twice about how we use AI in medicine. It's called "Large language models improve physician accuracy but lead to false reliance."
Jane: And Tom, that title is almost a warning, isn't it? It's saying these tools genuinely help doctors get more answers right, but there's a catch hiding in there. That "false reliance" part is the thing that keeps me up at night.
Tom: Exactly. And I've got Lu and Meng with us today to help unpack this. Lu, you've been working on medical AI for years now. What's your first reaction to that title?
Lu: My first reaction is relief, honestly. We've seen so many papers claiming AI will replace doctors or that AI is useless in the clinic. This one actually measures what happens when a real physician works with the system. And the headline finding is solid: doctors went from about seventy-one percent accuracy on their own to nearly eighty-three percent when they had the AI assistant helping them.
Meng: But that title is doing a lot of work. "Lead to false reliance" — that's not just a warning label. That's a specific failure mode they identified. And I want to know how they measured that, because that's the part that's going to determine whether this tool is safe to deploy.
Jane: And that's the thing that makes this paper different. It's not just "AI is good" or "AI is bad." It's "AI helps, but here's exactly when it hurts." And the hurt part is subtle. It's not about the AI being wrong in obvious ways. It's about the AI being wrong in ways that look convincing.
Tom: Right, and that's the false reliance. The doctors in this study were more likely to trust the AI when it showed them citations — actual source documents that supposedly supported its answer. And when those citations looked legitimate, doctors would abandon their own correct answers and follow the AI into a mistake.
Lu: And that's the scary part for me. The citations weren't always actually supporting the answer. The paper found that when a citation was judged as supportive, the AI's answer was correct about ninety-three percent of the time. But when no citation was supportive, the AI was still right about sixty-six percent of the time. So the signal is real, but it's not perfect. And the doctors were treating it as much more reliable than it actually was.
Meng: So the system is giving doctors a false sense of security. They see a citation, they think "this is grounded, this is verified," and they stop questioning. That's the automation bias we've seen in other fields, but it's especially dangerous in medicine.
Jane: And that's why this paper matters beyond just dermatology. It's showing us a pattern that's going to apply to any field where AI gives answers with sources attached. Law, engineering, finance — anywhere people are making high-stakes decisions based on AI output with citations.
Tom: So we've got a tool that genuinely improves performance, but it also creates a new kind of risk. And the question is, how do we build systems that keep the benefit without the danger? That's what we're going to dig into next.
Summary: Tom: So we've established that this paper, "Large language models improve physician accuracy but lead to false reliance," found real benefits but also real risks. Jane, walk us through the actual setup of the study, because I think the design is what makes the findings so compelling.
Jane: Sure. They built a system called CORA — Citation-Oriented Retrieval Assistant. It's basically an AI that doesn't just answer questions from memory. It goes and looks up medical guidelines, textbooks, and case reports, and then it gives you an answer with the actual sources it used. Think of it like a doctor who says "here's my diagnosis, and here's the textbook page that supports it."
Lu: And the clever part is how they tested it. They had forty-six dermatologists answer medical questions first without any help, then with CORA's answer and citations shown to them. Each doctor saw sixteen questions, so that's seven hundred thirty-six total decisions. And the doctors had to rate whether each citation actually supported the AI's answer before they could revise their own answer.
Meng: So the doctors weren't just passively accepting the AI. They were actively evaluating the evidence. And even with that active evaluation, they still fell into the trap. That's what makes this so concerning.
Tom: Right. So the doctors improved from about seventy-one percent to eighty-three percent accuracy overall. That's a big win. But when you break it down, you see the problem. When the AI was correct and the doctor was wrong, the doctor adopted the AI's answer about sixty-four percent of the time. That's good. But when the AI was wrong and the doctor was right, the doctor stuck with their own answer only about sixty-five percent of the time. That's not good enough.
Jane: And then it gets worse when you factor in the citations. When the citation was judged as supporting the AI's answer, doctors adopted correct advice seventy-seven percent of the time. But when the AI was wrong and the citation looked supportive, doctors resisted the bad advice only thirty-five percent of the time. So the citation made them more likely to follow the AI, whether the AI was right or wrong.
Lu: That's the grounding miscalibration. The doctors were using the citations as a proxy for correctness. And the citations were genuinely informative — they did correlate with the AI being right. But the correlation wasn't strong enough to justify the level of trust the doctors placed in them.
Meng: And here's the thing that really bothers me as an engineer. The system was built to be transparent. It shows you the sources. It lets you evaluate them. And that transparency actually made the problem worse in some cases, because it gave the doctors a false sense of verification.
Tom: So transparency alone isn't the solution. You can't just say "we showed them the sources, so it's their fault if they trusted them." The system needs to do more.
Jane: Exactly. And that's what we're going to talk about next — what the paper suggests we should actually do about this problem.
Improvements: Tom: So we've got this problem where citations make doctors trust the AI more, even when the AI is wrong. Jane, what does the paper suggest we actually do about this?
Jane: The paper is pretty clear that we need to change how we evaluate these systems. Accuracy alone isn't enough. You have to measure what happens when the system is wrong. They call it reliance-stratified reporting — you report performance separately for cases where the doctor and the AI agree, and cases where they disagree, and especially cases where the AI is wrong.
Lu: And that's a big shift. Most studies just report the overall accuracy number. But this paper shows that a system can raise average accuracy while making its errors harder to catch. The dangerous cases are the ones where the AI is wrong and the doctor was right — those are the ones where the system can actively harm patients.
Meng: So what's the actual fix? Are we talking about changing the interface, changing the training, changing the citations themselves?
Jane: The paper suggests a few directions. One is flagging when retrieval was weak — telling the doctor "we couldn't find strong evidence for this answer." Another is making the system distinguish between a source that's on-topic and a source that actually proves the answer. A citation can mention the right disease and the right drug without actually containing the evidence that makes the answer correct.
Lu: And that's a really important distinction. In the study, a lot of the citations were topically relevant — they mentioned the right condition, they were about the right subject. But they didn't actually support the specific claim the AI was making. And the doctors couldn't reliably tell the difference.
Tom: So the fix is to make the system itself better at evaluating its own evidence. Instead of just showing the doctor a source and saying "here you go," the system should say "this source supports this specific claim, and here's why."
Meng: That's a much harder engineering problem. It's not just retrieval and generation anymore. You need the system to actually reason about whether the evidence supports the conclusion. That's a different kind of capability.
Jane: And the paper also suggests something more radical — requiring an explicit check before a doctor can revise their answer based on the AI. Like, forcing them to articulate why they're changing their mind. That kind of cognitive forcing function has been shown to reduce over-reliance in other studies.
Lu: But I think the deeper point is that we need to design for the failure case. We know these systems will be wrong sometimes. The question is whether we can make the system's confidence signal align with its actual correctness. Right now, the citation support signal is correlated with correctness, but it's not reliable enough. The system needs to be honest about when it's uncertain.
Tom: So the improvements are about making the system more honest, not just more accurate. And that's a fundamentally different design philosophy.
Jane: Exactly. And that's what we're going to explore next — what the actual paper says in detail, starting from the beginning.
First Page: Tom: So we've talked about the big picture and the proposed fixes. Now let's actually look at the paper itself. "Large language models improve physician accuracy but lead to false reliance" — Jane, what does the first page tell us?
Jane: The first page sets up the core tension. Large language models are really good at medical tasks — they can process complex clinical information and generate relevant responses. But they also hallucinate, they can be outdated, and they don't show you where their answers come from. That's why the authors built CORA with retrieval-augmented generation.
Lu: And the key insight on that first page is that retrieval introduces a new failure mode. The generated answer might not actually reflect the retrieved sources, even when it has citations attached. So you can have a system that looks grounded — it shows you sources — but the sources don't actually support the answer.
Meng: So the citations are a form of trust signal, but they're not a reliable one. And the paper is saying that's a problem we haven't been paying enough attention to.
Jane: Right. And the first page also explains the structure of the paper. They did two things: first, they benchmarked CORA against regular LLMs on dermatology questions. Second, they ran the reader study with the forty-six physicians. The benchmark shows the system works. The reader study shows how doctors actually use it.
Tom: And the benchmark results are interesting. CORA was never worse than the base models — it was always at least as good, and often better. The gains were biggest for smaller models. The weakest model improved by about eight point six percentage points on the standard benchmark, and a whopping twenty-two point seven percentage points on the newer cases.
Lu: That's the part that really matters for real-world use. The cases published after the models' training cutoff — those are the ones where the AI can't just memorize the answer. It has to actually retrieve and reason over new information. And that's where retrieval-augmentation shines.
Meng: So the system is most valuable exactly when it's needed most — when the AI hasn't seen the information before. That's a strong argument for this approach.
Jane: And the first page also introduces the reader study design. Each physician answered questions unaided, then saw CORA's answer with citations, rated the citations, and could revise their answer. That design is what lets them measure the false reliance effect.
Tom: So the first page sets up the whole story: the problem, the system, the evaluation approach, and the key finding that citations create both benefit and risk. And that risk is what we need to keep thinking about.
Lu: And I think the most important line on that first page is the last one — it says the same signal that promotes appropriate reliance can also drive deference to incorrect advice. That's the fundamental tension we're dealing with.
Tom: And that tension is what makes this paper so important. It's not saying "don't use AI in medicine." It's saying "use it, but understand how it can go wrong, and design for that."
Conclusion: Tom: Alright, we've spent this whole episode on "Large language models improve physician accuracy but lead to false reliance." Jane, let's wrap this up for our listeners.
Jane: Let's do it. The core finding is that retrieval-augmented AI genuinely helps doctors — accuracy went from about seventy-one percent to eighty-three percent in their study. But the citations that make the system trustworthy also create a danger. When a citation looks supportive, doctors are more likely to follow the AI, even when the AI is wrong.
Lu: And the key mechanism is what they call grounding miscalibration. The citations are informative — they do correlate with correctness. But the correlation isn't strong enough. Doctors treat a supportive-looking citation as near-proof, when it's really just a hint.
Meng: For me, the practical takeaway is that we can't just add citations and call it safe. We need systems that flag weak retrieval, that distinguish between on-topic sources and sources that actually prove the claim, and that force doctors to think before they change their answers.
Tom: And the paper's broader message is that we need to evaluate these systems differently. Accuracy alone hides the danger. We need to look at what happens when the system is wrong, and whether the system's confidence signals align with reality.
Jane: That's the part that could change how we build and test medical AI everywhere. Not just dermatology — any field where AI gives answers with sources attached. The pattern is going to be the same.
Lu: And I think the most hopeful part is that the problem is identifiable and fixable. We know where the failure is. We know what causes it. We can design systems that are more honest about their uncertainty.
Tom: So that's "Large language models improve physician accuracy but lead to false reliance." A paper that gives us a real benefit and a real warning. Thanks to Lu and Meng for joining us, and thanks to all our listeners.
Jane: Next up, we've got a paper on something completely different — I'm told it's about using language models to design new materials. Should be fun.
Tom: Sounds great. Until then, keep questioning your sources, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language