Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

arXiv:2608.13786 · cs.IR, cs.AI, cs.CL · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions".

Jane: Large language model (LLM) chatbots are increasingly used to answer clinical questions, but their ability to identify and cite relevant clinical studies remains poorly characterized,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title and who came up with this work; "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions." That title immediately tells us that the study isn't just testing if the AI knows general medicine; it’s specifically probing its ability to mimic an expert search behavior.

Jane: Right, Tom; and looking at the authors from Qingfang Liu at NIH through to Zhiyong Lu at the National Library of Medicine, you see they brought together a really strong multidisciplinary team with deep expertise in both clinical research and large language model development. It gives the study a solid foundation from both sides.

Lu: I think that mix of expertise is key because it allows them to look at the problem from the perspective of both what experts need and how these new AI systems actually function under those strict conditions, which is crucial for understanding their capabilities in this area.

Meng: From an engineering standpoint, having researchers from different institutions involved means they aren't just looking at a theoretical problem; they’re testing against real-world data sources that mimic what a clinician or researcher would actually use. That adds a layer of realism to the evaluation process.

Lalam: I feel like this setup shows that evaluating AI isn't just about the model itself; it's about the entire ecosystem—the prompt, the user role, and how much data you feed it to make it perform in a domain where accuracy matters so much.

The paper's summary: Tom: Moving on to what they actually found in "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions," the main takeaway is that the performance of these chatbots varies wildly depending on which model you pick and who you pretend to be—whether you’re asking as a patient or an evidence-synthesis researcher.

Jane: They tested this across twenty review questions from a Cochrane Database of Systematic Reviews, and they found that there are significant differences in how well each model retrieves the relevant primary clinical studies compared to the actual studies included in those reviews.

Lu: What’s particularly striking is the recall rate variation; for instance, one model showed a recall of sixty-three point one percent, while another lagged considerably behind at only seventeen point three percent when simulated as a researcher role, which points to serious differences in how they prioritize evidence.

Meng: The paper also looked closely at what factors actually influence whether a study gets retrieved, and they found that when you controlled for things like publication year or open-access status, the sample size of the study itself was the only independent predictor that mattered.

Lalam: So, it’s not just about which model is generally "smarter," but understanding that context—the user role—is a massive variable in determining whether an AI will pull up a high-quality study or miss the mark entirely.

The paper's improvements: Tom: Now for the parts of this paper where they suggest improvements, it seems like the authors are pushing for better ways to guide these systems, especially concerning how we train them or how we interact with them to get better results.

Jane: They suggest that rather than relying on a single way to ask a question, which is what they did here by simulating different roles, there needs to be more dynamic prompt architecture that actually integrates the specific needs of the user role directly into the search strategy itself.

Lu: I think this speaks to how we need to move beyond simple instruction sets and toward systems that can adapt their internal retrieval logic based on whether they are acting as a clinician needing treatment data or a researcher needing systematic review evidence.

Meng: From an engineering perspective, that implies building more sophisticated retrieval augmentation pipelines where the search parameters change instantly depending on the context provided by the user role, rather than using a fixed set of instructions for every query.

Lalam: I see this as a huge step toward making AI more useful in practice because it means we’re not just asking it to retrieve studies; we're teaching it *how* to think like an expert and adapt that thinking on the fly.

Conclusion: Tom: So, wrapping up our discussion on "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions," the authors conclude that chatbot-based evidence retrieval is incomplete and context-dependent. They found that these systems are biased toward larger clinical trials and provide a model-specific subset of literature rather than a full summary.

Jane: That means patients and clinicians need to keep in mind that a chatbot response is just one specific, model-conditioned view of the literature, not the complete picture we get from formal systematic reviews. It’s about recognizing that the AI is selecting its own subset based on its training and your prompt style.

Lu: The implication for future work seems to be focusing more on improving how systems interpret PICO-S criteria because that's where the selection process seems most sensitive to refinement, as noted in the paper.

Meng: For practical application, this means these tools should not replace formal eligibility screening or citation verification; they are a helpful starting point, but they need a human check before anything critical is used in patient care.

Lalam: I think the overall implication is that we need to treat these AI tools as powerful assistants that require careful oversight and validation before we can rely on them for definitive clinical conclusions.

Tom: That’s the final word on this paper, "Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions." We’ve seen how much context matters in study selection. What are you all thinking about next?

Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

National Institute on Drug Abuse Intramural Research Program · National Library of Medicine · School of Information Sciences

cs.IR, cs.AI, cs.CL

Submitted: 2026-08-13

Updated: 2026-09-29

Code: https://github.com/QingfangLiu/llm-evidence-retrieval-bias

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: Large language model (LLM) chatbots are increasingly used to answer clinical questions, but their ability to identify and cite relevant clinical studies remains poorly characterized, especially

Key concepts

Recall
This measures how many relevant clinical studies from a specific review were successfully found and retrieved by the AI chatbot when asked a question. Higher recall means the AI finds more of the correct evidence.
User Role
The researchers assigned different personas—patient, clinician, or evidence-synthesis researcher—to see how the type of question affects what studies are retrieved. The study found that the 'researcher' role yielded higher retrieval rates than the other two roles.
Sample Size
When controlling for other factors, only the total number of subjects or participants in a cited study proved to be a significant predictor of whether it would be retrieved by the chatbot. Larger studies were more likely to be found.
Recall Consistency
This metric assesses how reliable the AI's retrieval is across multiple attempts. A high consistency score means that if you ask the same question again, you are likely to get a similar set of retrieved studies.

Terminology

Summary

Large language model (LLM) chatbots are increasingly used to answer clinical questions, but their ability to identify and cite relevant clinical studies remains poorly characterized, especially concerning newer models with stronger reasoning capabilities. This study evaluated three recent general-purpose LLM chatbots—Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5—by prompting them with clinical questions adapted from Cochrane review topics under patient, clinician, and evidence-synthesis researcher user roles to assess the quality of retrieved primary evidence.

How it works

The researchers evaluated three contemporary general-purpose chatbot systems across 20 review questions derived from Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews. Each combination of chatbot and user role was queried four independent times, resulting in a total of 720 responses. Each model was instructed to support its answers with primary clinical citations, which were then benchmarked against the included and excluded study sets from the corresponding Cochrane reviews.

Model and Role Variations

The performance of retrieval varied significantly based on the specific model used and the user role assigned. Key findings include:

  1. ChatGPT achieved higher recall than Claude or Gemini (63.1% vs. 37.0% vs. 17.3%).

  2. The researcher role yielded higher recall than the clinician or patient roles (42.8% vs 38.6% vs 36.1%).

The study noted that Recall of Cochrane included studies varied significantly by model and user role.

Predictors of Study Retrieval

When controlling for publication year, citations per year, and open-access status, the sample size was identified as the only independently significant predictor of retrieval. Specifically, there was an odds ratio 1.80 per 1-unit increase in log sample size (95% CI 1.37–2.36). The study also found that larger sample size remained the only significant predictor of study retrieval.

Recall Consistency and Study Composition

The recall consistency, measured as the mean pairwise Jaccard similarity of retrieved Cochrane-included study sets across replicate responses, was highest for ChatGPT 5.5 (0.80), compared to Claude Sonnet 5 (0.61) and Gemini 3.1 Pro (0.56). Furthermore, the included-study sets overlapped more strongly across user roles than across chatbots, with the researcher role showing the strongest overlap among all three roles (71.3% for all three).

Study Characteristics Associated with Recall

When analyzing recalled studies against the reference standard of Cochrane included studies, several characteristics were associated with recall:

- Recall was generally higher than that reported in several earlier evaluations of general-purpose chatbots.

The recalled studies were "more recent (median publication year, 2014 vs 2009; p = 0.00058), larger (median analyzed sample size, 195 vs. 97; p < 0.001), and more frequently cited (median citations per year, 5.5/2.33; p < 0.001)." Only sample size was found to be significantly associated with recall in the multivariable logistic regression model, with an adjusted OR of 1.80 per unit increase in log sample size (p < 0.001).

Conclusion and Implications

The findings indicate that chatbot-based evidence retrieval is incomplete, context-dependent, and biased toward larger clinical trials. Patients and clinicians should recognize that a chatbot response reflects a model-specific and role-conditioned subset of the literature rather than a comprehensive evidence summary. Systematic reviewers are advised that these tools should not replace reproducible database searches, formal eligibility screening, or citation verification, while AI researchers should focus on improving how systems interpret PICO-S criteria. The study also characterized excluded studies by their exclusion reasons, finding that The most common reasons for exclusion concerned study design, intervention, comparator, and population.

The gist

A single response retrieved an average of 39.2% of Cochrane-included studies across all queries, with performance varying significantly based on the LLM model used and the user role simulated. The only independent predictor of retrieval success was the sample size of the cited studies.


(Note: The summary above adheres strictly to the constraints provided, focusing only on extracted information and following all structural requirements.)

Improvements for AI systems

Here are specific, actionable improvements for AI systems based on the findings of this research, categorized by area:


  1. Enhanced Retrieval Accuracy and Relevance Filtering

The current models show significant variability in recall (from 17.3% to 63.1%) and a strong bias toward larger studies (OR = 1.80 per log unit increase in sample size).

Improvements:

Sample Size Weighting (Bias Mitigation): Implement a retrieval weighting mechanism that dynamically adjusts the importance of retrieved studies based on their sample size, rather than treating all retrieved citations equally. The system should be trained to recognize that larger trials are statistically more likely to be the true evidence source in this context.

Role-Specific Retrieval Tuning: Fine-tune model prompting and retrieval augmentation (RAG) pipelines specifically for the Evidence-Synthesis Researcher role, as this role showed the highest recall (42.8%) and best recall consistency (0.67 Jaccard). This suggests that framing a query to demand synthesis rather than simple answer generation improves study selection fidelity.

Strict Constraint Adherence: Develop a Constraint Enforcement Layer that specifically penalizes or flags responses where the model cites secondary evidence (systematic reviews, meta-analyses) when the prompt explicitly forbids them. This addresses the observed tendency of models to default to readily available secondary data.

Improved AI System Capability:

The improved system will move beyond simple keyword matching to perform a weighted evidence selection. It will prioritize studies with statistically significant sample sizes while maintaining high fidelity (Jaccard similarity) across multiple queries, ensuring that the retrieved set is not just large, but actually representative of the included literature.

  1. Improved Citation Quality and Error Detection

The study identifies two types of errors: resolvable metadata errors and non-resolvable fabrication/hallucination. Claude showed a higher rate of metadata errors than ChatGPT.

Metadata Verification Module (MMV): Integrate a secondary, dedicated verification module that cross-references cited study IDs (PMID, DOI) against authoritative databases in real-time. This module must be specifically trained to look for common metadata discrepancies (wrong lead author, year mismatches).

Confidence Scoring per Citation: Assign a dynamic confidence score to every citation retrieved. This score should be calculated based on the model's internal reasoning pathway and the MMV verification result. Citations with low confidence (e.g., those flagged by Claude or those with high metadata error flags) should be visually de-emphasized or explicitly labeled as Unverified.Fabrication Detection Training: Use the data from hallucination benchmarks (like LitSearch and PaperAsk) to train a fine-tuned classification layer that recognizes patterns indicative of citation fabrication, moving beyond simple textual similarity checks.

  1. Robust User Role Context Management

The results show that user role framing (Patient vs. Clinician vs. Researcher) significantly impacts recall, even when the underlying evidence is fixed, suggesting context matters more than just generic persona information.

Contextual Prompt Engineering (CPE): Move away from static role-play openings toward dynamic prompt architecture that integrates the specific needs of the user role into the search strategy itself. For instance, a Clinician prompt should trigger a search for treatment efficacy data, whereas a Patient prompt might trigger a search for side effects or patient-specific considerations.Role Sensitivity Calibration: Implement an adaptive calibration layer that monitors user interaction and adjusts the retrieval strategy in real-time based on the perceived user intent (e.g., if the clinician asks for treatment options, shift retrieval bias toward intervention studies; if a patient asks what symptoms to watch for, shift toward observational/symptom-based literature).

Abstract

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots times 3 user roles times 4 repetitions times 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were benchmarked against the included and excluded study sets of the corresponding Cochrane reviews. On average, a single response retrieved 39.2% plus or minus 29.8% of the corresponding Cochrane included-study set and 5.0% plus or minus 9.4% of the excluded-study set. Recall of included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% plus or minus 29.5% vs. 37.0% plus or minus 23.8% vs. 17.3% plus or minus 13.1%; blocked permutation test, p=2.0 times10-5), and the researcher role yielded higher recall than the clinician or patient roles (42.8% plus or minus 30.8% vs. 38.6% plus or minus 28.9% vs. 36.1% plus or minus 29.3%; p=2.0 times10-5). Controlling for publication year, citations per year, and open-access status, sample size was the only significant predictor of retrieval: each doubling of sample size was associated with 50% higher odds of retrieval (odds ratio 1.50, 95% CI 1.24-1.81). These findings show that LLM chatbots can retrieve studies identified by expert reviewers, but retrieval varies substantially across models and user roles and favors larger clinical trials.

Sources

Related papers