Large language models improve physician accuracy but lead to false reliance

arXiv:2608.00817 · cs.AI, cs.HC, stat.AP · Submitted 2026-08-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Large language models improve physician accuracy but lead to false reliance".

Jane: The paper was written by Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl et al. from German Cancer Research Center and University Heidelberg and University Medical Center Mannheim and Ruprecht-Karl University of Heidelberg and Medical University of Vienna and Pontificia Universidad Católica de Chile and University Medical Center Rostock and Heidelberg University Hospital.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's going to make you think twice about how we use AI in medicine. It's called "Large language models improve physician accuracy but lead to false reliance."

Jane: And Tom, that title is almost a warning, isn't it? It's saying these tools genuinely help doctors get more answers right, but there's a catch hiding in there. That "false reliance" part is the thing that keeps me up at night.

Tom: Exactly. And I've got Lu and Meng with us today to help unpack this. Lu, you've been working on medical AI for years now. What's your first reaction to that title?

Lu: My first reaction is relief, honestly. We've seen so many papers claiming AI will replace doctors or that AI is useless in the clinic. This one actually measures what happens when a real physician works with the system. And the headline finding is solid: doctors went from about seventy-one percent accuracy on their own to nearly eighty-three percent when they had the AI assistant helping them.

Meng: But that title is doing a lot of work. "Lead to false reliance" — that's not just a warning label. That's a specific failure mode they identified. And I want to know how they measured that, because that's the part that's going to determine whether this tool is safe to deploy.

Jane: And that's the thing that makes this paper different. It's not just "AI is good" or "AI is bad." It's "AI helps, but here's exactly when it hurts." And the hurt part is subtle. It's not about the AI being wrong in obvious ways. It's about the AI being wrong in ways that look convincing.

Tom: Right, and that's the false reliance. The doctors in this study were more likely to trust the AI when it showed them citations — actual source documents that supposedly supported its answer. And when those citations looked legitimate, doctors would abandon their own correct answers and follow the AI into a mistake.

Lu: And that's the scary part for me. The citations weren't always actually supporting the answer. The paper found that when a citation was judged as supportive, the AI's answer was correct about ninety-three percent of the time. But when no citation was supportive, the AI was still right about sixty-six percent of the time. So the signal is real, but it's not perfect. And the doctors were treating it as much more reliable than it actually was.

Meng: So the system is giving doctors a false sense of security. They see a citation, they think "this is grounded, this is verified," and they stop questioning. That's the automation bias we've seen in other fields, but it's especially dangerous in medicine.

Jane: And that's why this paper matters beyond just dermatology. It's showing us a pattern that's going to apply to any field where AI gives answers with sources attached. Law, engineering, finance — anywhere people are making high-stakes decisions based on AI output with citations.

Tom: So we've got a tool that genuinely improves performance, but it also creates a new kind of risk. And the question is, how do we build systems that keep the benefit without the danger? That's what we're going to dig into next.

Summary: Tom: So we've established that this paper, "Large language models improve physician accuracy but lead to false reliance," found real benefits but also real risks. Jane, walk us through the actual setup of the study, because I think the design is what makes the findings so compelling.

Jane: Sure. They built a system called CORA — Citation-Oriented Retrieval Assistant. It's basically an AI that doesn't just answer questions from memory. It goes and looks up medical guidelines, textbooks, and case reports, and then it gives you an answer with the actual sources it used. Think of it like a doctor who says "here's my diagnosis, and here's the textbook page that supports it."

Lu: And the clever part is how they tested it. They had forty-six dermatologists answer medical questions first without any help, then with CORA's answer and citations shown to them. Each doctor saw sixteen questions, so that's seven hundred thirty-six total decisions. And the doctors had to rate whether each citation actually supported the AI's answer before they could revise their own answer.

Meng: So the doctors weren't just passively accepting the AI. They were actively evaluating the evidence. And even with that active evaluation, they still fell into the trap. That's what makes this so concerning.

Tom: Right. So the doctors improved from about seventy-one percent to eighty-three percent accuracy overall. That's a big win. But when you break it down, you see the problem. When the AI was correct and the doctor was wrong, the doctor adopted the AI's answer about sixty-four percent of the time. That's good. But when the AI was wrong and the doctor was right, the doctor stuck with their own answer only about sixty-five percent of the time. That's not good enough.

Jane: And then it gets worse when you factor in the citations. When the citation was judged as supporting the AI's answer, doctors adopted correct advice seventy-seven percent of the time. But when the AI was wrong and the citation looked supportive, doctors resisted the bad advice only thirty-five percent of the time. So the citation made them more likely to follow the AI, whether the AI was right or wrong.

Lu: That's the grounding miscalibration. The doctors were using the citations as a proxy for correctness. And the citations were genuinely informative — they did correlate with the AI being right. But the correlation wasn't strong enough to justify the level of trust the doctors placed in them.

Meng: And here's the thing that really bothers me as an engineer. The system was built to be transparent. It shows you the sources. It lets you evaluate them. And that transparency actually made the problem worse in some cases, because it gave the doctors a false sense of verification.

Tom: So transparency alone isn't the solution. You can't just say "we showed them the sources, so it's their fault if they trusted them." The system needs to do more.

Jane: Exactly. And that's what we're going to talk about next — what the paper suggests we should actually do about this problem.

Improvements: Tom: So we've got this problem where citations make doctors trust the AI more, even when the AI is wrong. Jane, what does the paper suggest we actually do about this?

Jane: The paper is pretty clear that we need to change how we evaluate these systems. Accuracy alone isn't enough. You have to measure what happens when the system is wrong. They call it reliance-stratified reporting — you report performance separately for cases where the doctor and the AI agree, and cases where they disagree, and especially cases where the AI is wrong.

Lu: And that's a big shift. Most studies just report the overall accuracy number. But this paper shows that a system can raise average accuracy while making its errors harder to catch. The dangerous cases are the ones where the AI is wrong and the doctor was right — those are the ones where the system can actively harm patients.

Meng: So what's the actual fix? Are we talking about changing the interface, changing the training, changing the citations themselves?

Jane: The paper suggests a few directions. One is flagging when retrieval was weak — telling the doctor "we couldn't find strong evidence for this answer." Another is making the system distinguish between a source that's on-topic and a source that actually proves the answer. A citation can mention the right disease and the right drug without actually containing the evidence that makes the answer correct.

Lu: And that's a really important distinction. In the study, a lot of the citations were topically relevant — they mentioned the right condition, they were about the right subject. But they didn't actually support the specific claim the AI was making. And the doctors couldn't reliably tell the difference.

Tom: So the fix is to make the system itself better at evaluating its own evidence. Instead of just showing the doctor a source and saying "here you go," the system should say "this source supports this specific claim, and here's why."

Meng: That's a much harder engineering problem. It's not just retrieval and generation anymore. You need the system to actually reason about whether the evidence supports the conclusion. That's a different kind of capability.

Jane: And the paper also suggests something more radical — requiring an explicit check before a doctor can revise their answer based on the AI. Like, forcing them to articulate why they're changing their mind. That kind of cognitive forcing function has been shown to reduce over-reliance in other studies.

Lu: But I think the deeper point is that we need to design for the failure case. We know these systems will be wrong sometimes. The question is whether we can make the system's confidence signal align with its actual correctness. Right now, the citation support signal is correlated with correctness, but it's not reliable enough. The system needs to be honest about when it's uncertain.

Tom: So the improvements are about making the system more honest, not just more accurate. And that's a fundamentally different design philosophy.

Jane: Exactly. And that's what we're going to explore next — what the actual paper says in detail, starting from the beginning.

First Page: Tom: So we've talked about the big picture and the proposed fixes. Now let's actually look at the paper itself. "Large language models improve physician accuracy but lead to false reliance" — Jane, what does the first page tell us?

Jane: The first page sets up the core tension. Large language models are really good at medical tasks — they can process complex clinical information and generate relevant responses. But they also hallucinate, they can be outdated, and they don't show you where their answers come from. That's why the authors built CORA with retrieval-augmented generation.

Lu: And the key insight on that first page is that retrieval introduces a new failure mode. The generated answer might not actually reflect the retrieved sources, even when it has citations attached. So you can have a system that looks grounded — it shows you sources — but the sources don't actually support the answer.

Meng: So the citations are a form of trust signal, but they're not a reliable one. And the paper is saying that's a problem we haven't been paying enough attention to.

Jane: Right. And the first page also explains the structure of the paper. They did two things: first, they benchmarked CORA against regular LLMs on dermatology questions. Second, they ran the reader study with the forty-six physicians. The benchmark shows the system works. The reader study shows how doctors actually use it.

Tom: And the benchmark results are interesting. CORA was never worse than the base models — it was always at least as good, and often better. The gains were biggest for smaller models. The weakest model improved by about eight point six percentage points on the standard benchmark, and a whopping twenty-two point seven percentage points on the newer cases.

Lu: That's the part that really matters for real-world use. The cases published after the models' training cutoff — those are the ones where the AI can't just memorize the answer. It has to actually retrieve and reason over new information. And that's where retrieval-augmentation shines.

Meng: So the system is most valuable exactly when it's needed most — when the AI hasn't seen the information before. That's a strong argument for this approach.

Jane: And the first page also introduces the reader study design. Each physician answered questions unaided, then saw CORA's answer with citations, rated the citations, and could revise their answer. That design is what lets them measure the false reliance effect.

Tom: So the first page sets up the whole story: the problem, the system, the evaluation approach, and the key finding that citations create both benefit and risk. And that risk is what we need to keep thinking about.

Lu: And I think the most important line on that first page is the last one — it says the same signal that promotes appropriate reliance can also drive deference to incorrect advice. That's the fundamental tension we're dealing with.

Tom: And that tension is what makes this paper so important. It's not saying "don't use AI in medicine." It's saying "use it, but understand how it can go wrong, and design for that."

Conclusion: Tom: Alright, we've spent this whole episode on "Large language models improve physician accuracy but lead to false reliance." Jane, let's wrap this up for our listeners.

Jane: Let's do it. The core finding is that retrieval-augmented AI genuinely helps doctors — accuracy went from about seventy-one percent to eighty-three percent in their study. But the citations that make the system trustworthy also create a danger. When a citation looks supportive, doctors are more likely to follow the AI, even when the AI is wrong.

Lu: And the key mechanism is what they call grounding miscalibration. The citations are informative — they do correlate with correctness. But the correlation isn't strong enough. Doctors treat a supportive-looking citation as near-proof, when it's really just a hint.

Meng: For me, the practical takeaway is that we can't just add citations and call it safe. We need systems that flag weak retrieval, that distinguish between on-topic sources and sources that actually prove the claim, and that force doctors to think before they change their answers.

Tom: And the paper's broader message is that we need to evaluate these systems differently. Accuracy alone hides the danger. We need to look at what happens when the system is wrong, and whether the system's confidence signals align with reality.

Jane: That's the part that could change how we build and test medical AI everywhere. Not just dermatology — any field where AI gives answers with sources attached. The pattern is going to be the same.

Lu: And I think the most hopeful part is that the problem is identifiable and fixable. We know where the failure is. We know what causes it. We can design systems that are more honest about their uncertainty.

Tom: So that's "Large language models improve physician accuracy but lead to false reliance." A paper that gives us a real benefit and a real warning. Thanks to Lu and Meng for joining us, and thanks to all our listeners.

Jane: Next up, we've got a paper on something completely different — I'm told it's about using language models to design new materials. Should be fun.

Tom: Sounds great. Until then, keep questioning your sources, everyone.

Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Titus J. Brinker

German Cancer Research Center · University Heidelberg · University Medical Center Mannheim · Ruprecht-Karl University of Heidelberg · Medical University of Vienna · Pontificia Universidad Católica de Chile · University Medical Center Rostock · Heidelberg University Hospital

cs.AI, cs.HC, stat.AP

Submitted: 2026-08-01

Updated: 2026-08-11

Code: https://github.com/DBO-DKFZ/CORA

License: http://creativecommons.org/publicdomain/zero/1.0/

Importance score: 65/100

The gist: The paper presents CORA (Citation-Oriented Retrieval Assistant), an agentic retrieval-augmented generation (RAG) system for dermatological decision support, and evaluates both its standalone

Key concepts

False Reliance
This occurs when doctors trust an AI's answer too much because the AI provides a citation, even if that citation does not actually support the answer. This leads doctors to follow incorrect advice.
Grounding Miscalibration
This is a problem where the signal provided by citations is correlated with correctness but is not strong enough to justify high levels of trust. Doctors treat supportive-looking citations as near-proof, even when they are only hints.
Reliance-Stratified Reporting
This proposed method for evaluating AI systems involves reporting performance separately based on whether the doctor and AI agree or disagree, especially focusing on cases where the AI is wrong to identify dangerous failures.

Terminology

Summary

The paper presents CORA (Citation-Oriented Retrieval Assistant), an agentic retrieval-augmented generation (RAG) system for dermatological decision support, and evaluates both its standalone performance and its effects on physician decision-making. The authors state: We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making.

System design: CORA uses a knowledge base comprising 72 condition-specific clinical practice guidelines from the European Academy of Dermatology and Venereology (EADV), an authoritative dermatology reference work: Fitzpatrick's Dermatology in General Medicine (9th edition), and a collection of 7,936 dermatology case reports published between 2004 and 2024. Retrieval proceeds as a bounded multi-step loop orchestrated by a Qwen3 model (Qwen3-235B-A22B-Instruct-2507), with documents embedded using Snowflake Arctic Embed V2 and reranked with Mixedbread Rerank Large v1. The loop terminates on a judgement of sufficiency or on reaching the iteration cap of three retrieval rounds.

Benchmark evaluation: The authors evaluated CORA against five backbone LLMs (GPT-5, Llama-4 Scout, Mistral Large 2, Qwen 2.5, Gemma-3) on two datasets. The first, DermBenchQA, comprised 4,855 single-best-answer questions drawn from four medical question-answering (QA) benchmarks. The second, DermCaseQA, comprised 998 open-ended questions generated from 893 PubMed-indexed dermatology case reports published in 2025 or later, after the training cutoff of every backbone model evaluated. On the 2,207 matched DermBenchQA items, CORA maintained performance to the corresponding baseline across all models (non-inferiority tests, all corrected P-values ≤ 0.004), with gains scaling inversely with baseline capability: GPT-5 gained 0.3 pp (92.7% to 93.0%), while Gemma 3 gained 8.6 pp (76.9% to 85.5%). On the contamination-resistant DermCaseQA set (437 matched items), retrieval gains were larger and more uniform: Gemma 3 improved by 22.7 pp (37.5% to 60.2%), Qwen 2.5 by 18 pp (46.5% to 64.5%), Llama 4 by 10.9 pp (51.3% to 62.2%), Mistral Large 2 by 9.9 pp (55.1% to 65.0%), and GPT-5 by 3.4 pp (77.1% to 80.5%).

Reader study: The authors conducted "a preregistered within-subjects reader study in which 46 physicians answered dermatology multiple-choice questions unaided and then with CORA assistance, rating the support of each cited passage before being permitted to revise their answer. Participants were 46 physicians in dermatology from 21 countries, including 3 dermatology residents. Each physician answered 16 questions from a 320-item subsample of DermBenchQA, yielding 736 paired decisions. CORA's answers were generated by Llama-4 Scout. The study found that mean accuracy increased from 70.8% (95% CI 67.0%, 74.5%) without assistance to 82.6% (95% CI 80.0%, 85.2%) after participants viewed CORA's answer and cited sources (Wilcoxon signed-rank test, P<0.001, n=46 physicians). Specifically, assistance corrected 110 initially incorrect responses and overturned 23 initially correct responses into incorrect; 498 responses remained correct and 105 remained incorrect."

Citation support findings: Across the 192 questions rated by three physicians, at least one citation was judged supportive for 80.2% (citation support rate; 95% bootstrap CI: 74.5%, 85.4%), while 60.6% of individual cited sources were rated as supportive (citation precision; 95% CI: 56.2%, 64.9%). The presence of supporting citations was strongly associated with correctness: "CORA answered 92.9% of questions correctly when at least one citation supported its answer, compared with 65.8% when no citation was supportive (odds ratio (OR), 6.76; 95% CI: 2.73, 16.77; P<0.001). At the physician-decision level, Final answers were correct in 87.7% of responses when at least one citation was judged to support the CORA answer, compared with 65.5% when no citation was judged supportive (OR, 3.75; 95% CI: 2.48, 5.67; P<0.001, n=736 physician-question decisions). Among initially incorrect physician answers, physicians reached the correct final answer in 64.6% of responses when at least one citation supported the CORA answer, compared with 23.9% when no supporting citation was present (OR, 5.79; 95% CI: 2.82, 11.90; P<0.001; n=215 physician-question decisions)."

Reliance analysis: Using the Schemmer et al. appropriate-reliance framework, the authors report: "Physicians adopted correct CORA advice in 64.3% of cases where they were wrong (RAIR; 95% CI: 56.9%, 71.4%) and resisted incorrect CORA advice in 64.6% of cases where they were right (RSR; 95% CI: 51.0%, 78.4%), placing them in the appropriate-reliance quadrant. However, stratifying by perceived citation support revealed a directional dissociation between the two forms of reliance": "When CORA was correct, RAIR was higher among decisions in which at least one citation was judged supportive (93/121; 76.9%; 95% CI: 67.8%, 84.9%) than among those in which no citation was judged supportive (17/50; 34.0%; 95% CI: 20.8%, 47.7%; 42.9 pp increase). When CORA was incorrect, the association reversed: RSR was lower with perceived citation support (8/23; 34.8%; 95% CI: 17.2%, 52.6%) than without it (23/25; 92.0%; 95% CI: 79.2%, 100%; 57.2 pp decrease). The authors note: when physicians judged that a citation supported an incorrect CORA answer, they abandoned an initially correct response in 15 of 23 cases."

Key conclusion: The authors term the observed pattern grounding miscalibration: perceived evidentiary support that is not reliably aligned with the correctness of the answer it accompanies. They conclude: "The failure is not in physicians' judgement of citations but in the fact that the display of citations made an incorrect answer look grounded. A citation that is topically relevant but does not establish the generated claim may thus be more dangerous than an unsupported answer, because it manufactures an appearance of verifiability that can lead a physician to abandon a correct assessment. They further state: These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk."

Improvements for AI systems

Based on this paper, here are the specific improvements I can implement in AI systems:

Improvement: Add a post-generation verification layer that checks whether cited sources actually support the generated answer's specific claims, not just topical relevance.

What the improved system can do:

  • Before presenting citations to a user, run a claim-level entailment check: extract each factual claim from the answer and verify it against the cited source text

  • Flag citations that are topically relevant (same disease/drug) but do not contain the specific evidence supporting the answer

  • Display a support confidence score per citation (e.g., This source mentions the condition but does not confirm the specific treatment recommendation)

  • When support confidence is low, add a warning: The cited source may not fully support this answer. Please verify independently.

Improvement: Implement a two-sided reliability signal that explicitly distinguishes evidence supports the answer from evidence contradicts the answer.

Improvement: Replace single accuracy metrics with a two-dimensional evaluation that separates adoption of correct advice (RAIR) from resistance to incorrect advice (RSR).

Improvement: Prioritize retrieval for cases published after the model's training cutoff, where the paper showed the largest gains (up to 22.7 pp for weaker models).

Improvement: Design the user interface to make citation support assessment more deliberate and less automatic.

Improvement: Make the retrieval sufficiency decision visible to the user.

Improvement: The paper showed weaker models (Gemma 3) gained 22.7 pp with retrieval while strong models (GPT-5) gained only 3.4 pp. Implement adaptive retrieval intensity based on model capability.

Improvement: The paper showed physicians varied widely in their reliance patterns (37 improved, 8 unchanged, 1 declined). Implement personalized reliance calibration.

Improvement: The paper found that citations can be topically relevant but not claim-supporting. Modify the generation process to explicitly check for contradiction.

Improvement: The paper's DermCaseQA set showed retrieval gains were larger on post-cutoff data. Implement contamination detection.

Abstract

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

Related papers