ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
cs.CL
Submitted: 2026-03-02
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings.
Terminology
Abstract
Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce ClinConsensus, a Chinese medical benchmark of 2, 500 expert-curated cases spanning 36 specialties, 12 task themes, multiple difficulty levels, and lay-facing versus professional-facing settings. Each case is paired with 30 case-specific binary rubric criteria. To evaluate whether responses satisfy enough physician-authored criteria, we propose Clinician-Anchored Coverage Score (CACS), a physician-calibrated threshold metric instantiated at k=10, and develop a dual-judge framework combining a GPT-5.1 grader with a physician-supervised Qwen3-8B judge. Evaluating 11 frontier LLMs, we find a persistent coverage gap: Rubric Accuracy ranges from 39.6% to 52.1%, whereas CACS@10 ranges from 17.8% to 32.9%, leaving a 19.2--21.9 point gap across models. Stratified analyses further reveal substantial variation across reasoning, evidence use, structured extraction, medication instructions, follow-up, and dialogue register. These results suggest that medical LLM evaluation should measure thresholded, rubric-grounded clinical coverage rather than average partial correctness.
Sources
- Large Language Models in Healthcare
- MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
- Measuring Massive Multitask Language Understanding
- Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
- Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
- Large Language Models as Agents in the Clinic
- Polaris: A Safety-focused LLM Constellation Architecture for Healthcare
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
- LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering