Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
summary
The gist
Large language model (LLM) chatbots are increasingly used to answer clinical questions, but their ability to identify and cite relevant clinical studies remains poorly characterized, especially
In short
Researchers tested three advanced AI chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and GPT-5.5) answering clinical questions based on Cochrane review topics using different user roles (patient, clinician, researcher). Findings showed ChatGPT had the best recall, but retrieval success was primarily predicted by the sample size of the cited studies rather than other factors.
Key concepts
- Recall
- This measures how many relevant clinical studies from a specific review were successfully found and retrieved by the AI chatbot when asked a question. Higher recall means the AI finds more of the correct evidence.
- User Role
- The researchers assigned different personas—patient, clinician, or evidence-synthesis researcher—to see how the type of question affects what studies are retrieved. The study found that the 'researcher' role yielded higher retrieval rates than the other two roles.
- Sample Size
- When controlling for other factors, only the total number of subjects or participants in a cited study proved to be a significant predictor of whether it would be retrieved by the chatbot. Larger studies were more likely to be found.
- Recall Consistency
- This metric assesses how reliable the AI's retrieval is across multiple attempts. A high consistency score means that if you ask the same question again, you are likely to get a similar set of retrieved studies.
Terminology used across episodes
This episode discusses
- Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions · Paper Radio
- Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution · Paper Radio
- Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study
The paper
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions · Read on arXiv
Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu
National Institute on Drug Abuse Intramural Research Program · National Library of Medicine · School of Information Sciences
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions".
Jane: Large language model (LLM) chatbots are increasingly used to answer clinical questions, but their ability to identify and cite relevant clinical studies remains poorly characterized,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the title and who came up with this work; "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions." That title immediately tells us that the study isn't just testing if the AI knows general medicine; it’s specifically probing its ability to mimic an expert search behavior.
Jane: Right, Tom; and looking at the authors from Qingfang Liu at NIH through to Zhiyong Lu at the National Library of Medicine, you see they brought together a really strong multidisciplinary team with deep expertise in both clinical research and large language model development. It gives the study a solid foundation from both sides.
Lu: I think that mix of expertise is key because it allows them to look at the problem from the perspective of both what experts need and how these new AI systems actually function under those strict conditions, which is crucial for understanding their capabilities in this area.
Meng: From an engineering standpoint, having researchers from different institutions involved means they aren't just looking at a theoretical problem; they’re testing against real-world data sources that mimic what a clinician or researcher would actually use. That adds a layer of realism to the evaluation process.
Lalam: I feel like this setup shows that evaluating AI isn't just about the model itself; it's about the entire ecosystem—the prompt, the user role, and how much data you feed it to make it perform in a domain where accuracy matters so much.
The paper's summary: Tom: Moving on to what they actually found in "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions," the main takeaway is that the performance of these chatbots varies wildly depending on which model you pick and who you pretend to be—whether you’re asking as a patient or an evidence-synthesis researcher.
Jane: They tested this across twenty review questions from a Cochrane Database of Systematic Reviews, and they found that there are significant differences in how well each model retrieves the relevant primary clinical studies compared to the actual studies included in those reviews.
Lu: What’s particularly striking is the recall rate variation; for instance, one model showed a recall of sixty-three point one percent, while another lagged considerably behind at only seventeen point three percent when simulated as a researcher role, which points to serious differences in how they prioritize evidence.
Meng: The paper also looked closely at what factors actually influence whether a study gets retrieved, and they found that when you controlled for things like publication year or open-access status, the sample size of the study itself was the only independent predictor that mattered.
Lalam: So, it’s not just about which model is generally "smarter," but understanding that context—the user role—is a massive variable in determining whether an AI will pull up a high-quality study or miss the mark entirely.
The paper's improvements: Tom: Now for the parts of this paper where they suggest improvements, it seems like the authors are pushing for better ways to guide these systems, especially concerning how we train them or how we interact with them to get better results.
Jane: They suggest that rather than relying on a single way to ask a question, which is what they did here by simulating different roles, there needs to be more dynamic prompt architecture that actually integrates the specific needs of the user role directly into the search strategy itself.
Lu: I think this speaks to how we need to move beyond simple instruction sets and toward systems that can adapt their internal retrieval logic based on whether they are acting as a clinician needing treatment data or a researcher needing systematic review evidence.
Meng: From an engineering perspective, that implies building more sophisticated retrieval augmentation pipelines where the search parameters change instantly depending on the context provided by the user role, rather than using a fixed set of instructions for every query.
Lalam: I see this as a huge step toward making AI more useful in practice because it means we’re not just asking it to retrieve studies; we're teaching it *how* to think like an expert and adapt that thinking on the fly.
Conclusion: Tom: So, wrapping up our discussion on "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions," the authors conclude that chatbot-based evidence retrieval is incomplete and context-dependent. They found that these systems are biased toward larger clinical trials and provide a model-specific subset of literature rather than a full summary.
Jane: That means patients and clinicians need to keep in mind that a chatbot response is just one specific, model-conditioned view of the literature, not the complete picture we get from formal systematic reviews. It’s about recognizing that the AI is selecting its own subset based on its training and your prompt style.
Lu: The implication for future work seems to be focusing more on improving how systems interpret PICO-S criteria because that's where the selection process seems most sensitive to refinement, as noted in the paper.
Meng: For practical application, this means these tools should not replace formal eligibility screening or citation verification; they are a helpful starting point, but they need a human check before anything critical is used in patient care.
Lalam: I think the overall implication is that we need to treat these AI tools as powerful assistants that require careful oversight and validation before we can rely on them for definitive clinical conclusions.
Tom: That’s the final word on this paper, "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions." We’ve seen how much context matters in study selection. What are you all thinking about next?
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck