DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models".
Jane: The paper was written by Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie et al. from Shanghai Jiao Tong University and SII and SPIRAL Lab and Generative AI Research Lab (GAIR) and Shanghai Chest Hospital and Beijing Anzhen Hospital, Capital Medical University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re diving into a paper that’s been making waves in the medical AI world: "DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models." Jane, this one’s from a team at Shanghai Jiao Tong University, and honestly, the title alone tells you they’re serious about testing these models where it actually matters.
Jane: Absolutely, Tom. And I love the name "DiagnosisArena" because it really captures what they’ve built — a proving ground, a place where AI models have to step into the ring and face real clinical cases. It’s not a multiple-choice test from a textbook. It’s like putting a medical student in front of a patient and saying, "Okay, what’s going on here?"
Tom: Right, and that’s the big shift. We’ve seen models ace medical exams, but this paper argues those exams have become too easy. They’re saturated. So the team went and pulled over a thousand real case reports from top journals like NEJM and The Lancet, and they structured them into a format that mimics actual clinical encounters.
Jane: And that’s the key, Tom. They didn’t just grab random cases. They built a whole pipeline to make sure these cases are genuinely hard — they filtered out anything a model could solve with simple recall. They even had doctors and AI experts double-check each case to make sure there was enough information to reach a diagnosis, but not so much that it becomes obvious.
Lu: If I can jump in here — I’m Lu, by the way. What excites me most is that they’re not just testing knowledge. They’re testing reasoning. The cases require you to connect symptoms, physical exam findings, and test results into a coherent diagnosis. That’s the kind of cognitive work we hope AI can eventually support in clinics.
Meng: And as an engineer, I’m curious about the scale. They ended up with one thousand one hundred thirteen cases across twenty-eight specialties. That’s a solid benchmark. But what really caught my eye is the performance gap — even the best model, o3, only got about fifty-one percent accuracy. That’s barely better than a coin flip.
Tom: Yeah, that number is sobering. And it’s not like they used weak models. They tested o1, DeepSeek-R1, GPT-4o, Claude, all the big names. And the results show that when you take away the multiple-choice crutch, these models struggle. Jane, what do you think that says about where we are with AI in medicine?
Jane: I think it says we’re still in the early innings. The models are great at recalling facts, but they’re not yet great at the messy, integrative thinking that diagnosis requires. And that’s exactly why a benchmark like this is so valuable — it gives us a clear, honest picture of what’s missing.
Lu: And that’s the exciting part. Once you can measure the gap, you can start closing it. This paper gives the research community a target to aim for.
Meng: Yeah, but I also wonder — is the gap because the models lack medical knowledge, or because they lack a way to structure their thinking? That’s the question I’d love to dig into next.
Jane: Great question, Meng. And that’s exactly where we’re heading in the next segment — we’re going to look at how they actually built this benchmark and what makes it so much harder than what came before.
Summary of the Paper: Tom: Welcome back. We’re still on "DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models." Jane, let’s get into the meat of it — how did they actually build this thing?
Jane: So the process was really meticulous, Tom. They started with over four thousand case reports from ten top-tier journals. Then they used a mix of rule-based filtering and AI models to restructure each case into four clean sections: case information, physical exam, diagnostic tests, and the final diagnosis. That last part is the ground truth they use to score the models.
Lu: And here’s the clever part — they didn’t just take all those cases. They ran an iterative filtering process. They used models like GPT-4o and DeepSeek-V3 to try each case, and if any of them could solve it easily, they threw it out. That way, they kept only the cases that genuinely require deep reasoning.
Meng: That’s a smart way to avoid the "too easy" problem. But they also did something even more interesting — they had AI experts and board-certified physicians review each remaining case. If the diagnosis was ambiguous or the case didn’t have enough information, it got cut. So what’s left is a set of cases where a correct answer is actually possible, but only if you reason well.
Tom: And that’s why the benchmark is so tough. It’s not just about having the right knowledge — it’s about applying it in a messy, real-world context. The paper shows that even the best reasoning models, like o1 and DeepSeek-R1, score in the seventeen percent to thirty-one percent range. That’s a huge drop from the ninety percent plus we see on traditional medical benchmarks.
Jane: And that drop is the whole point, right? It shows that the old benchmarks were measuring something different — mostly recall. DiagnosisArena measures something harder: the ability to synthesize information and make a judgment call under uncertainty.
Lu: Exactly, Jane. And I think that’s why this paper is so important. It’s not just a new dataset. It’s a new way of thinking about evaluation. We’re moving from "can the model remember the answer" to "can the model reason its way to the answer."
Meng: But I want to push back a little. They also created a multiple-choice version of the same cases, and the models did much better on that — o1 jumped to almost sixty-two percent. Doesn’t that suggest the models actually know the answers, but just can’t express them in an open-ended format?
Jane: That’s a really good point, Meng. And the paper addresses that directly. They argue that multiple-choice questions give the model a huge hint — it just has to pick the best option, not generate a diagnosis from scratch. That’s like giving a student a list of possible answers and saying "one of these is right." It’s a fundamentally easier task.
Tom: So the open-ended format is the real test. And the fact that models struggle with it tells us something important about their limitations. They can recognize the right answer when they see it, but they can’t always produce it on their own.
Lu: And that’s the gap we need to close. The next question is — can we actually improve these models, or is there a fundamental limit? That’s what we’re going to talk about next.
Improvements Suggested: Tom: Back on the air with "DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models." We’ve talked about the benchmark itself, but now let’s get into what the paper suggests we do about these results. Jane, what’s the path forward?
Jane: Well, Tom, the paper doesn’t propose a specific new model, but it does point to a clear direction. The biggest takeaway is that reasoning ability matters more than raw knowledge. The models that did best, like o3 and o1, are specifically trained to think step-by-step. That’s why they outperformed models like GPT-4o, which is powerful but not designed for deep reasoning.
Lu: And I think that’s the key insight. The paper shows that when you give a model time to "think" — to generate a long chain of reasoning before answering — it does much better. That’s the whole idea behind inference-time scaling. It’s not about making the model bigger; it’s about letting it think longer.
Meng: But that comes at a cost. More thinking means more compute, which means more money and more latency. In a real clinical setting, you can’t always wait five minutes for a diagnosis. So there’s a trade-off between accuracy and practicality.
Tom: That’s a real concern, Meng. But the paper also suggests another angle — that multiple-choice formats are a crutch. If we want models to be useful in real clinics, we need to train them to generate diagnoses from scratch, not just pick from a list.
Jane: And that’s where the benchmark can help. By giving researchers a hard, realistic test, DiagnosisArena can drive the development of better reasoning strategies. Maybe it’s better training data, maybe it’s new prompting techniques, maybe it’s models that can ask for more tests when they’re uncertain.
Lu: I’d love to see that last one. Imagine a model that says, "I’m not sure, but if we ran this specific test, it would help me narrow it down." That’s how real doctors work. That’s the kind of behavior we should be encouraging.
Meng: And that would require a different kind of benchmark too — one that rewards asking the right questions, not just giving the right answer. But that’s a bigger project.
Tom: Definitely. But for now, DiagnosisArena gives us a solid foundation. It’s a hard test, and it shows us where we stand. And the fact that even the best models are struggling means there’s a lot of room for improvement.
Jane: And that’s exciting, Tom. It means the field is still wide open. The next breakthrough could come from anywhere — a new training method, a new model architecture, or a new way of thinking about clinical reasoning.
Lu: I think the biggest impact will be on how we design AI systems for healthcare. Instead of just building bigger models, we’ll need to build systems that can reason, ask questions, and collaborate with doctors. That’s a much richer vision.
Meng: And it’s a vision that’s grounded in reality. This paper gives us the tools to measure progress toward that vision. That’s what I appreciate most.
Tom: Well said. And that brings us to our final segment, where we’ll wrap up and look at the bigger picture.
Conclusion: Tom: And we’re back for the final stretch. We’ve been talking about "DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models," and I have to say, this paper has given us a lot to think about.
Jane: It really has, Tom. The core message is simple but powerful: current AI models are not ready for real-world clinical diagnosis. They can ace exams, but they struggle when faced with the messy, complex cases that doctors deal with every day. And that’s a crucial reality check.
Lu: And it’s a constructive one, too. By building a benchmark that’s actually hard, the authors have given the research community a clear target. We now know exactly what we need to improve — not just knowledge, but reasoning.
Meng: I’d add that the paper also highlights the importance of evaluation design. The fact that models do so much better on multiple-choice questions shows that the format of the test can dramatically change the results. That’s a lesson that goes beyond medicine.
Tom: That’s a great point, Meng. And it makes me think about the broader implications. If we’re going to use AI in high-stakes fields like healthcare, we need to be honest about its limitations. Benchmarks like this help us do that.
Jane: And they also help us track progress. A year from now, we can run the same models on DiagnosisArena and see if they’ve gotten better. That’s how we know we’re moving in the right direction.
Lu: I’m optimistic. The gap between fifty-one percent and one hundred percent is large, but it’s not insurmountable. With better reasoning models, better training data, and better evaluation methods, we can close it. And when we do, AI could become a truly valuable partner for doctors.
Meng: And that’s the dream, right? Not replacing doctors, but supporting them — helping them see things they might have missed, and giving them more time to focus on their patients.
Tom: Well said, Meng. And on that note, we’re going to wrap up our discussion of "DiagnosisArena." It’s a paper that’s both humbling and inspiring. Humbling because it shows how far we still have to go, and inspiring because it gives us a clear path forward.
Jane: Thanks for listening, everyone. We’ll be back next time with another paper from the arXiv. Until then, keep thinking, keep questioning, and keep pushing the boundaries of what’s possible.
Tom: Take care, and see you on the next episode.
Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
Shanghai Jiao Tong University · SII · SPIRAL Lab · Generative AI Research Lab (GAIR) · Shanghai Chest Hospital · Beijing Anzhen Hospital, Capital Medical University
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
Code: https://github.com/SPIRAL-MED/DiagnosisArena
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 64/100
Key concepts
- DiagnosisArena
- A new benchmark designed to test LLMs on real clinical cases, moving beyond simple textbook exams. It requires models to synthesize symptoms, physical findings, and test results into a coherent diagnosis.
- Diagnostic Reasoning
- The ability AI models demonstrate by connecting various pieces of information—symptoms, tests, etc.—to arrive at a correct medical conclusion. The paper shows models struggle with this complex cognitive work.
- Open-ended vs. Multiple-choice
- The difficulty of the test format. Models perform poorly when required to generate a diagnosis from scratch (open-ended) but show much higher accuracy when given a list of possible answers (multiple-choice).
Terminology
Summary
Summary
This paper introduces DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess the professional-level diagnostic competence of large language models (LLMs) in complex clinical scenarios. The authors state: "Given the limitations of existing medical benchmarks in evaluating advanced diagnostic reasoning, we present DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess professional-level diagnostic competence."
The benchmark consists of 1,113 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, derived from clinical case reports published in 10 top-tier medical journals, including Lancet, NEJM, JAMA, and others. The authors emphasize that "the inclusion of highly challenging cases and a multi-stage curation process culminates in a professional-grade benchmark comprising 1,113 structured clinical cases across 28 medical specialties, designed specifically for evaluating the diagnostic reasoning capabilities of LLMs in complex clinical scenarios."
The construction pipeline involves four stages: data collection, data structuring, iterative filtering, and expert-AI collaborative verification. Initially, 4,175 case reports were collected from top-tier journals. The authors note: we conducted an extensive review of numerous medical journals and ultimately selected 10 target journals as our data sources, from which we collected a total of 4,175 case reports.
Data structuring converts raw reports into standardized Markdown format with four sections: case information, physical examination, diagnostic tests, and final diagnosis. Iterative filtering uses models like Baichuan-M1, DeepSeek-V3, and GPT-4o to eliminate overly simple cases, retaining 1,783 cases. Expert-AI collaborative verification then excludes cases where consensus cannot be reached, resulting in the final 1,113 cases.
The evaluation methodology uses GPT-4o as a judge to categorize diagnostic outputs as identical
, relevant
, or irrelevant
, with only identical
considered correct. Models generate five candidate diagnoses, and both top-1 and top-5 accuracy are calculated. The authors also created a multiple-choice version, DiagnosisArena-MCQ, for comparison.
Key experimental findings reveal that even the most advanced reasoning models, o3, o1, and DeepSeek-R1, achieve only 51.12%, 31.09%, and 17.79% accuracy, respectively.
Other models struggle to surpass 20% accuracy. The authors highlight that Models endowed with explicit reasoning capabilities demonstrate a clear advantage in clinical diagnostic tasks,
noting that reasoning-enhanced models like QwQ-32B achieve 25.69% accuracy, and DeepSeek-R1 shows a 13.66% improvement over its base model DeepSeek-V3.
The paper also examines data leakage effects by analyzing model performance across publication years, concluding that instances of data leakage within our benchmark are exceedingly rare, as indicated by the absence of significant divergence in model performance for cases published before versus after the training data cut-off date.
A significant finding is the performance gap between open-ended and multiple-choice formats: "By converting the diagnostic task in DiagnosisArena into the multiple-choice format based on model-generated diagnoses as DiagnosisArena-MCQ, we observe a marked increase in model performance, with o1 reaching 61.90%, further suggesting that multiple-choice formats inherently reduce task difficulty and thus fail to accurately reflect the models' true abilities in addressing complex clinical problems."
The case study analysis reveals that current models struggle with complex diagnostic reasoning, as they tend to prioritize the likelihood of common diseases rather than inferring based on available clues.
The authors conclude that current models have not yet fully adapted to the complex reasoning demands of clinical diagnostic scenarios,
and they hope DiagnosisArena will contribute to advancing reasoning abilities in the field of medicine.
Improvements for AI systems
Based on the paper’s findings, here are specific improvements to AI systems and the resulting capabilities:
1. Implement a “Clue-Fidelity” Reasoning Constraint
-
Improvement: Modify the model’s decoding or fine-tuning objective to penalize reasoning paths that ignore or contradict explicit patient-specific clues (e.g., “highly mobile element in LVOT,” “no anatomical relationship with mitral/aortic valve”). Add a verification step that forces the model to explicitly map each clue to a candidate diagnosis or state why it’s irrelevant before finalizing the answer.
-
Capability: The AI will avoid the failure mode shown in Figure 5, where DeepSeek-R1 overlooked AMVT despite multiple supporting imaging findings. It will produce diagnoses that are traceable to the case data, reducing reliance on common-disease priors.
2. Add a “Differential-Expansion” Module for Rare/Structural Conditions
-
Improvement: After generating initial top-5 diagnoses, run a secondary retrieval-augmented generation (RAG) pass over a curated database of rare structural cardiac/anatomical anomalies (e.g., AMVT, accessory chordae, subaortic membranes). Use the model’s own uncertainty (e.g., low confidence in top-1) to trigger this expansion, and re-rank the combined list.
-
Capability: The system will include rare but correct diagnoses in its top-5, improving Top-5 accuracy from the observed 17–51% range toward higher recall, especially for cases where imaging shows unusual mobile structures.
3. Train a “Clinical-Reasoning-Contrast” Loss
-
Improvement: Fine-tune models on pairs of cases: one where the correct diagnosis is rare (e.g., AMVT) and one where a common mimic (e.g., papillary fibroelastoma) is present. Use a contrastive loss that explicitly rewards the model for distinguishing subtle imaging differences (e.g., “filamentous” vs. “solid,” “attachment to posterior LVOT wall” vs. “valvular”). Include negative samples from the paper’s case study.
-
Capability: The AI will better discriminate between structurally similar but diagnostically distinct conditions, reducing the 13.66% gap seen between DeepSeek-V3 and DeepSeek-R1 by teaching the model to reason from evidence rather than pattern-matching.
4. Build a “Leakage-Resilient Evaluation” Mode
-
Improvement: For any new case, compute the publication date and compare model performance against a rolling baseline (as in Figure 4). If accuracy on post-cutoff cases drops significantly, automatically flag the model as “knowledge-memorizing” rather than reasoning. Use this as a gating metric before clinical deployment.
-
Capability: The system will provide a reliable readiness score for real-world use, ensuring that high performance on DiagnosisArena is not an artifact of training-data leakage, and alerting developers to models that fail on genuinely novel cases.
5. Implement a “Multi-Stage Diagnostic Verification” Pipeline
-
Improvement: After the model outputs its top-5, run a second LLM (e.g., GPT-4o) as a judge to check if any output is “relevant” (score 1) but not “identical” (score 2). If so, prompt the original model to generate a more specific sub-diagnosis (e.g., from “cardiac mass” to “accessory mitral valve tissue”). Iterate up to 3 times.
-
Capability: The AI will convert vague but partially correct answers into precise final diagnoses, improving Top-1 accuracy beyond the current 51.12% ceiling, and better matching the granularity required in clinical handoffs.
6. Add a “Clue-Sufficiency” Pre-Checker
-
Improvement: Before answering, the model must classify the case as “sufficient,” “ambiguous,” or “insufficient” based on whether the provided physical exam and diagnostic tests uniquely support a diagnosis. If ambiguous, the model should explicitly list missing tests (e.g., “need TEE with 3D reconstruction to rule out AMVT”) rather than guessing.
-
Capability: The system will avoid overconfident wrong answers, and in real clinical use, it will recommend additional diagnostic steps—aligning with how physicians handle uncertainty, and reducing the risk of misdiagnosis from incomplete data.
7. Optimize for Open-Ended Output Over MCQ
-
Improvement: Given the paper’s finding that MCQ inflates performance (o1: 31% → 62%), fine-tune models specifically on open-ended diagnostic generation tasks, using DiagnosisArena’s open-ended format as the primary training signal. Disable any internal multiple-choice shortcuts during inference.
-
Capability: The AI will maintain high accuracy in real-world settings where options are not provided, avoiding the false sense of competence that MCQ-based benchmarks create.
Sources
- Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- An Empirical Study on Eliciting and Improving R1-like Reasoning Models
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- Measuring Massive Multitask Language Understanding
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
- O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
- MedS$^3$: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
- Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
- PubMedQA: A Dataset for Biomedical Research Question Answering
- Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
- From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
- O1 Replication Journey: A Strategic Progress Report -- Part 1
- Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Qwen2.5 Technical Report
- Baichuan-M1: Pushing the Medical Capability of Large Language Models
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- CMB: A Comprehensive Medical Benchmark in Chinese
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering