MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

arXiv:2609.12851 · cs.AI · Submitted 2026-09-11 · Read on arXiv

cs.AI

Submitted: 2026-09-11

Updated: 2026-09-11

License: http://creativecommons.org/licenses/by/4.0/

The gist: Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations.

Terminology

Abstract

Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.

Related papers