ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation".
Jane: The paper was written by Xiaoman Zhang, Hong-Yu Zhou, Xiaoli Yang, Oishi Banerjee, Julián N. Acosta et al. from Harvard Medical School and Gradient Health and Duke University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a very straightforward title: "ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation." And honestly, that title tells you exactly what it does, which I appreciate.
Jane: It really does, Tom. And I love that because in this field, sometimes you get these flashy names that hide what's actually going on. But ReXrank is basically saying, hey, we built a scoreboard for AI systems that write radiology reports, and we're putting it online for everyone to see.
Tom: A scoreboard. That's the perfect way to put it. So instead of every research group testing their model on their own private data and reporting their own numbers, ReXrank is saying, no, we all play on the same field now.
Jane: Exactly. And that's such a big deal because right now, if you read ten papers about chest X-ray report generation, you'll see ten different ways of evaluating. Some use one dataset, some use another, some report different metrics. It's really hard to know which model is actually better.
Lu: If I can jump in here, Jane, that's precisely the pain point this paper addresses. As someone who's worked on medical imaging models, I can tell you that comparing two systems right now is nearly impossible. You'd have to re-run everything yourself, and even then, the data splits might not match.
Tom: So Lu, you're saying this is a real headache in the field right now?
Lu: Absolutely. And it's not just a minor inconvenience. It means progress is slower because we can't tell what's actually working. If model A reports a zero point five on some metric and model B reports a zero point six, but they used different test sets, that comparison is meaningless.
Jane: And that's why the word "public" in the title matters so much. It's not just a benchmark that sits in a drawer somewhere. It's a living leaderboard where new models can be submitted and compared against everything that came before.
Tom: So it's like the ImageNet of radiology reports, but with a constantly updated scoreboard.
Jane: That's a great analogy, Tom. And the authors are from Harvard Medical School, which gives it some real credibility. They've also got collaborators from Gradient Health, which is the company that provided the private dataset.
Meng: I have to ask, though, as the engineer in the room, what does this actually mean for someone who wants to deploy one of these models in a hospital? Does a high score on ReXrank mean it's ready for clinical use?
Lu: That's a really important question, Meng. And the honest answer is, not necessarily. But it's a necessary first step. You need to know which models are worth even considering for clinical trials.
Tom: And that's the hook for what we're going to dig into next. Because the paper doesn't just build a leaderboard, it also tells us a lot about how these models actually perform across different datasets. But first, let's just appreciate what a public standard like this means for the field.
Summary: Tom: So we've established that ReXrank is a public leaderboard for AI radiology report generation. But what does the paper actually show us? What's the headline finding?
Jane: Well, Tom, the most striking result is that MedVersa, a model from Harvard, comes out on top. It achieves the best scores on three out of the four datasets they tested. And it consistently beats GPT-4V, which is OpenAI's flagship vision model.
Tom: That's a big deal, right? Because GPT-4V is this general-purpose model that can do almost anything. And here's a specialized medical model beating it on radiology reports.
Jane: Exactly. And it's not even close. On the private ReXGradient dataset, MedVersa gets a one/RadCliQ-v1 score of zero point nine eight, while GPT-4V only manages zero point six six. That's a huge gap.
Lu: And what's interesting to me, from a research perspective, is that this shows specialization still matters. Generalist models are impressive, but when you have a model trained specifically on medical data, it can outperform them significantly.
Meng: But hold on, I want to understand what these numbers actually mean. What is RadCliQ-v1? And why are they taking the reciprocal of it?
Jane: Good question, Meng. RadCliQ is a composite metric that combines several other metrics into one number. It's designed to be a single score that captures overall report quality. And the original metric is lower-is-better, meaning a score of zero would be perfect. So they take the reciprocal, one divided by the score, so that higher numbers mean better performance, which is more intuitive.
Tom: And it's not just one metric. The paper uses eight different metrics total, from simple ones like BLEU-two which measures word overlap, to more sophisticated ones like GREEN and FineRadScore, which use large language models to evaluate clinical accuracy.
Lu: That's actually one of the strengths of this work. They're not relying on a single number. They're giving you a full picture. A model might do well on lexical similarity but poorly on clinical accuracy, and you need to see both.
Meng: So it's like looking at a car's specs. You don't just look at horsepower. You look at fuel efficiency, safety ratings, reliability, all of it.
Jane: Perfect analogy, Meng. And the paper also shows something interesting about the datasets themselves. The IU X-ray dataset, which is relatively small and simple, gives very high scores across all models. But CheXpert Plus, which has a different distribution, shows much lower performance and higher variance.
Tom: So some datasets are just easier than others?
Jane: Right. And that's actually a really important insight. If you only test on IU X-ray, you might think your model is amazing. But when you test it on more challenging, diverse data, the performance drops significantly.
Lu: And that's why the private ReXGradient dataset is so valuable. It's ten thousand studies from sixty-seven different medical sites across the US. That's real-world diversity. And interestingly, the models show very low variance on this dataset, which suggests it's a reliable benchmark.
Tom: So we've got the headline results. MedVersa wins, specialized models beat generalists, and dataset choice matters a lot. But there's more to this paper. Next, we need to talk about how they actually built this thing and what improvements they're suggesting for the field.
Improvements: Tom: So we've covered what ReXrank is and what the results show. But what does this paper actually suggest we should do differently? What improvements is it pushing for?
Jane: The biggest one, Tom, is standardization. The paper is essentially arguing that the field needs to stop doing ad-hoc evaluations and start using a common framework. And they're not just talking about it, they built it.
Lu: And I think that's the key contribution here. They're saying, look, we've all been using MIMIC-CXR as a benchmark, but everyone uses different splits, different metrics, different preprocessing. It's chaos. ReXrank is an attempt to bring order to that chaos.
Meng: So it's like when software development moved from everyone having their own build system to standardized CI/CD pipelines. It just makes everything more reliable and comparable.
Jane: Exactly, Meng. And they also introduce this new dataset, ReXGradient, which is private. That's actually a really clever move because it prevents models from being overfit to public benchmarks. If you can't see the test data, you can't tune your model to it.
Tom: That's a really important point. Because with public datasets, there's always this risk that models are implicitly trained on them, even if the authors don't intend it. A private dataset is a true test of generalization.
Lu: And the results support that. The models show remarkably consistent performance on ReXGradient, with very tight confidence intervals. That suggests it's a high-quality, reliable benchmark that can really differentiate between models.
Meng: But I want to push back a little here. The paper also shows that models perform better when tested on the same distribution they were trained on. Like VLCI IU does great on IU X-ray but poorly on MIMIC-CXR. Doesn't that mean the leaderboard is just measuring how well a model fits a particular dataset?
Jane: That's a fair concern, Meng. But I think the paper addresses it by including multiple datasets with different distributions. A model that only does well on one dataset will rank lower overall than a model that generalizes across all of them.
Lu: And that's actually the right approach. You want a model that works in the real world, where you'll encounter all kinds of variations in imaging equipment, patient populations, and clinical practices. Testing on a single dataset doesn't tell you that.
Tom: So the improvement here is really about moving from single-dataset evaluations to multi-dataset, standardized evaluations with a mix of public and private data.
Jane: Right. And they also separate models into two categories: those that generate only findings and those that generate both findings and impressions. Because those are actually different tasks with different challenges.
Meng: That's a thoughtful design choice. It's like comparing a sprinter and a marathon runner. They're both runners, but you wouldn't put them in the same race.
Tom: And there's one more thing I want to get to. The paper also shows that models trained on multiple datasets tend to outperform those trained on a single dataset. That's a really practical insight for anyone building these systems. But let's save that for our next segment, where we'll look at the first page of the paper in detail.
First Page: Tom: We're back, and we're going to look at the actual first page of the ReXrank paper. Because sometimes the abstract and introduction tell you more than the whole rest of the paper.
Jane: And the first thing that jumps out, Tom, is the scale of this thing. They've got sixteen models from ten different institutions. That's not a small pilot study. That's a serious, comprehensive benchmark.
Lu: And they're using eight different metrics. That's a lot. Most papers use two or three. But here, they're trying to capture every dimension of report quality, from basic text similarity to clinical accuracy.
Meng: I noticed something else on that first page. They mention that the private ReXGradient dataset has ten thousand studies from sixty-seven medical sites. That's a massive amount of diverse data. Where did that come from?
Jane: It's from Gradient Health, which is one of the collaborating organizations. And the paper notes that the authors from Gradient Health are founders of the company. So there's a potential conflict of interest there, but they disclose it clearly.
Tom: That's good to see. Transparency about these things is important, especially in medical AI where trust is crucial.
Lu: And the first page also gives us a nice visual overview of the whole system. There's a figure showing the datasets, the models, the metrics, and the ranking process. It's a really clean way to understand what they've built.
Meng: I also noticed that they mention some models can only generate findings, while others can generate both findings and impressions. That distinction is important because impressions are the summary that referring physicians often read first.
Jane: Exactly. And the paper treats them as separate tasks, which makes sense. Generating a good impression requires synthesizing the findings into a concise clinical assessment. That's harder than just describing what you see.
Tom: So on that first page, we get the scope, the collaborators, the structure, and the key design decisions. It's a really well-organized introduction.
Lu: And I think the most important sentence on that page is where they say they're providing a standardized evaluation framework. That's the core contribution. Everything else, the datasets, the metrics, the leaderboard, is in service of that goal.
Meng: But I want to go back to something from earlier. The paper says MedVersa beats GPT-4V consistently. Is that really a fair comparison? GPT-4V is a general model, not specifically trained for radiology.
Jane: That's a fair point, Meng. But I think it's actually the point of the comparison. It shows that if you want a model for radiology report generation, you should use a specialized model, not a general-purpose one. That's a useful finding for anyone making deployment decisions.
Tom: And it also shows that the leaderboard can surface these kinds of insights. Without a standardized comparison, you might assume GPT-4V is the best choice because it's so capable in other domains.
Lu: Right. And that's the value of ReXrank. It gives you evidence-based guidance on which models to consider for clinical applications.
Conclusion: Tom: So we've spent this whole episode on ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation. Let's bring it all together.
Jane: The core idea, Tom, is that this paper provides a standardized way to evaluate AI models that generate radiology reports. It uses four datasets, eight metrics, and sixteen models to create a public leaderboard where anyone can submit their model and see how it ranks.
Lu: And the key findings are that MedVersa leads the pack, specialized medical models beat general-purpose ones like GPT-4V, and the choice of dataset dramatically affects measured performance.
Meng: For me, the most practical takeaway is that if you're building one of these systems, you should train on multiple datasets and test on diverse data. The paper shows that models trained on a single dataset don't generalize well.
Tom: And the bigger picture, the thing that gets me excited, is that this kind of standardization could accelerate the whole field. When everyone uses the same benchmark, progress becomes measurable and comparable.
Jane: That's right. And it's not just about radiology. The framework they've built could be extended to other medical imaging domains. The paper even says that explicitly, that it sets the stage for evaluation across the full spectrum of medical imaging.
Lu: I think that's the most exciting implication. This is a template for how to build trustworthy benchmarks in medical AI. If we apply this approach to pathology, dermatology, ophthalmology, we could see similar acceleration in those fields.
Meng: And from a deployment perspective, having a reliable leaderboard means hospitals and health systems can make more informed decisions about which AI tools to adopt. That's a real-world impact.
Tom: So ReXrank isn't just a scoreboard. It's a foundation for better, safer, more reliable medical AI.
Jane: Well said, Tom. And with that, we're going to wrap up our discussion of this paper. Thanks to Lu and Meng for joining us today.
Lu: Thanks for having me. This was a great conversation.
Meng: Yeah, really enjoyed it.
Tom: And to our listeners, if you want to check out the leaderboard yourself, it's at rexrank.ai. Next up, we've got a paper on something completely different, so stay tuned.
Jane: Until then, keep questioning, keep learning, and we'll see you on the next episode.
Xiaoman Zhang, Hong-Yu Zhou, Xiaoli Yang, Oishi Banerjee, Julián N. Acosta, Mohammed Baharoon, Josh Miller, Ouwen Huang, Pranav Rajpurkar
Harvard Medical School · Gradient Health · Duke University
cs.CV, cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/rajpurkarlab/CXR-Report-Metric
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 73/100
Key concepts
- ReXrank
- A public leaderboard created to standardize the evaluation of AI systems that generate radiology reports. It allows researchers to compare different models on a common field, moving away from private, inconsistent testing methods.
- MedVersa
- A model developed by Harvard that emerged as the top performer on the ReXrank leaderboard. It demonstrated superior performance compared to general-purpose models like GPT-4V when tested on medical radiology reports.
- RadCliQ-v1
- A composite metric used in the evaluation, combining several other metrics into a single score for overall report quality. The original metric is lower-is-better, so the authors use the reciprocal to make higher numbers indicate better performance.
- Standardization in AI Evaluation
- The paper argues that the field needs a common framework for evaluating models instead of ad-hoc testing. ReXrank aims to bring order by providing a consistent set of datasets and metrics for comparison.
Terminology
Summary
Summary
This paper introduces ReXrank (https://rexrank.ai), a public leaderboard and challenge designed to standardize the evaluation of AI-powered radiology report generation from chest X-ray images. The authors state: there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank... a public leaderboard and challenge for assessing AI-powered radiology report generation.
The framework incorporates ReXGradient, described as the largest test dataset consisting of 10,000 studies,
along with three public datasets: MIMIC-CXR, IU-Xray, and CheXpert Plus. ReXGradient is a private dataset provided by Gradient Health, which consists of 10,000 studies collected from 7,004 patients across 67 medical sites in the United States.
For public datasets, the authors utilize the official test splits of MIMIC-CXR (2,347 studies) and IU-Xray (590 studies), along with CheXpert Plus's validation set (200 studies) as no test split is available.
ReXrank employs 8 evaluation metrics: BLEU-2, BERTScore, SembScore, RadGraph-F1, RadCliQ-v1, RaTEScore, GREEN, and FineRadScore. The authors note: We default use RadCliQ-v1 as the primary metric.
For consistency, they present reciprocals of RadCliQ-v1 and FineRadScore so that higher values indicate better performance.
The leaderboard currently includes 16 models from 10 institutions: BiomedGPT IU, CheXagent, CheXpertPlus CheX, CheXpertPlus CheX MIMIC, CheXpertPlus MIMIC, Cvt2distilgpt2 IU, Cvt2distilgpt2 MIMIC, GPT4V, LLM-CXR, MAIRA-2, MedVersa, RadFM, RaDialog, RGRG, VLCI IU, and VLCI MIMIC. The framework separately assesses models capable of generating only findings sections and those providing both findings and impressions sections.
Key results show that MedVersa emerges as one of the top-performing models... with best 1/RadCliQ-v1 scores of 0.98 ± 0.05 on ReXGradient and 0.92 ± 0.02 on MIMIC-CXR.
However, its performance on the CheXpert Plus dataset is comparatively lower, ranking fourth with a 1/RadCliQ-v1 score of 0.72 ± 0.10 on the Findings.
MedVersa consistently outperforms GPT4V, the state-of-the-art generalist vision-language model, across multiple metrics and datasets.
Dataset analysis reveals that "IU X-ray stands out as the least challenging, consistently yielding high performance across models. In contrast, CheXpert Plus exhibits the highest variance and lower overall performance, likely due to its distinct data distribution and small validation set (200 studies). The private ReXGradient dataset
demonstrates remarkably low-performance variance across models, underscoring its high data quality and utility as a benchmark for assessing model robustness."
Additional findings include: "Models trained on multiple datasets (e.g., CheXpertPlus CheX MIMIC) tend to outperform those trained on individual datasets, suggesting that a multi-dataset training approach helps bridge the distributional gap and enhances generalization. Models perform better on training-distribution data, as
VLCI IU achieves superior performance on IU X-ray (RadCliQ-v1: 1.38 ± 0.04) compared to VLCI MIMIC (0.91 ± 0.04), while VLCI MIMIC performs better on MIMIC-CXR (0.68 ± 0.02 vs 0.60 ± 0.02 for VLCI IU)."
The comparison between findings-only and findings+impression tasks shows that "On ReXGradient, MedVersa shows slight performance degradation when generating both findings and impressions (1/RadCliQ-v1 decreasing from 1.01 ± 0.01 to 0.98 ± 0.05), while CheXpertPlus CheX MIMIC shows improvement (from 0.83 ± 0.01 to 0.85 ± 0.01). The authors attribute this to CheXpertPlus CheX MIMIC using
separate models for findings and impression generation, while MedVersa only uses a single model architecture."
The paper concludes that "By providing this standardized evaluation framework, ReXrank enables meaningful comparisons of model performance and offers crucial insights into their robustness across diverse clinical settings. Beyond its current focus on chest X-rays, ReXrank's framework sets the stage for comprehensive evaluation of automated reporting across the full spectrum of medical imaging."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems for radiology report generation:
-
Train models on combined datasets (MIMIC-CXR + CheXpert Plus + IU-Xray) rather than single datasets
-
Evidence: CheXpertPlus CheX MIMIC outperforms single-dataset models (0.85 vs 0.76 on ReXGradient)
-
Result: Better generalization across diverse clinical settings and patient populations
-
Use dedicated models for findings vs. impression sections instead of one unified model
-
Evidence: CheXpertPlus CheX MIMIC improved from 0.83 to 0.85 when using separate models
-
Result: More specialized and accurate section generation
-
Fine-tune models specifically to optimize RadCliQ-v1 (composite of BLEU, BERTScore, SembScore, RadGraph-F1)
-
Evidence: MedVersa achieves best 1/RadCliQ-v1 scores (1.01 on ReXGradient, 1.10 on MIMIC-CXR)
-
Result: Balanced improvement across all sub-metrics rather than gaming single metrics
-
Add ReXGradient-style diverse data (67 U.S. medical sites) during training
-
Evidence: Models show minimal variance (±0.01) on ReXGradient, indicating reliability
-
Result: More consistent performance across unseen institutions
-
Implement explicit handling of frontal + lateral image pairs
-
Evidence: Models like MAIRA-2 that accept technique info perform better on multi-view studies
-
Result: Better interpretation of complete radiological context
-
Generate clinically accurate reports with 1/RadCliQ-v1 scores above 1.0 on MIMIC-CXR (vs. current best of 1.10)
-
Maintain performance across 4+ different hospital systems with variance under 0.02
-
Produce both findings and impressions with section-specific optimization
-
Handle single and multi-view chest X-rays with appropriate contextual awareness
-
Achieve 0.98+ 1/RadCliQ-v1 on private datasets with 10,000 studies from 67 sites
-
Outperform GPT-4V by 30-50% on clinical metrics (RadGraph, RaTEScore, GREEN)
-
Provide reliable confidence intervals (±0.01) for deployment in clinical settings
-
Architecture: Use Swinv2 encoder + BERT decoder (proven by CheXpertPlus models)
-
Training: Sequential training for findings → impressions with shared visual features
-
Loss Function: Weighted combination of RadCliQ components (BLEU: 0.2, BERTScore: 0.2, SembScore: 0.3, RadGraph: 0.3)
-
Evaluation: Always report on ReXGradient + MIMIC-CXR + IU-Xray + CheXpert Plus for fair comparison
-
Data Augmentation: Include CheXpert Plus validation set (200 studies) for distribution-shift testing
Abstract
AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank, https://rexrank.ai, a public leaderboard and challenge for assessing AI-powered radiology report generation. Our framework incorporates ReXGradient, the largest test dataset consisting of 10,000 studies, and three public datasets (MIMIC-CXR, IU-Xray, CheXpert Plus) for report generation assessment. ReXrank employs 8 evaluation metrics and separately assesses models capable of generating only findings sections and those providing both findings and impressions sections. By providing this standardized evaluation framework, ReXrank enables meaningful comparisons of model performance and offers crucial insights into their robustness across diverse clinical settings. Beyond its current focus on chest X-rays, ReXrank's framework sets the stage for comprehensive evaluation of automated reporting across the full spectrum of medical imaging.
Sources
- MAIRA-2: Grounded Radiology Report Generation
- CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
- Cross-Modal Causal Intervention for Medical Report Generation
- Generating Radiology Reports via Memory-driven Transformer
- A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
- FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores
- RadGraph: Extracting Clinical Entities and Relations from Radiology Reports
- LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
- GREEN: Generative Radiology Report Evaluation and Error Notation
- RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
- CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT
- Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data
- Chest ImaGenome Dataset for Clinical Reasoning
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- BERTScore: Evaluating Text Generation with BERT
- MedVersa: A Generalist Foundation Model for Medical Image Interpretation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models