RadFusion: Towards Threshold-Controllable Radiology Report Generation

arXiv:2608.10505 · cs.AI, cs.CL, cs.CV · Submitted 2026-08-11 · Read on arXiv

Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz

Microsoft

cs.AI, cs.CL, cs.CV

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 95/100

The gist: RadFusion: Towards Threshold-Controllable Radiology Report Generation Summary This paper introduces RadFusion, a framework for threshold-controllable radiology report generation.

Terminology

Summary

RadFusion: Towards Threshold-Controllable Radiology Report Generation

Summary

This paper introduces RadFusion, a framework for threshold-controllable radiology report generation. The core problem addressed is that existing automated radiology report generation models, unlike perception models (image classifiers), offer no control over the sensitivity–specificity trade-off of their diagnostic content. This control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report cannot adapt to these scenarios or support ROC-based validation widely expected for regulatory clearance.

RadFusion equips report generation with threshold controllability by fusing three components:

  1. Perception Model (Classifier): A multi-label classifier (based on MedImageInsight) that provides per-disease confidence scores for 14 CheXpert classes. Given a threshold τ, it yields binary predictions, partitioning classes into positive (disease present) and negative (disease absent).

  2. Report Generation Model (QRad Auto-VQA): A VQA-based report generator that produces detailed, open-vocabulary radiology reports describing findings with attributes like location, severity, and progression. Its Auto-VQA reframing allows for follow-up queries to recover omitted information.

  3. LLM Rewriter: An off-the-shelf large language model (e.g., GPT-5) that rewrites the report so that its stated diagnoses align with the classifier's thresholded decisions while remaining grounded in the generator's descriptions. The rewriting instruction explicitly defines how to handle positive and negative classes, how to incorporate evidence, and how to maintain non-class content.

The key mechanism is that by sweeping the threshold τ from 0 to 1, RadFusion produces a family of reports spanning the full sensitivity–specificity spectrum. The evaluation uses a closed-loop protocol that converts the rewritten reports back to class labels and compares them against ground truth to trace an ROC curve.

Key Results on MIMIC-CXR:

  • Conformance to Classifier ROC: The threshold-controlled reports closely conform to the classifier's ROC curves. The paper states: the threshold-controlled reports closely conform to the classifier’s ROC curves, with the report ROC (red) closely tracking the classifier ROC (blue) across all 13 disease classes. The concave shape confirms a monotonic sensitivity–specificity trade-off.

  • Improved Factual Correctness: Combining the two model types improves diagnostic accuracy over uncontrolled generation. Specifically, sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. The ROC curves of threshold-controlled reports are above the non-controlled reports (green crosses).

  • Clinical Utility: The framework makes the sensitivity–specificity trade-off explicit and tunable, replacing the fixed, opaque operating point of conventional models. Low thresholds prioritize sensitivity (screening/triage); high thresholds prioritize specificity (confirmatory assessment).

Ablation Studies:

  • Alternative Classifier Implementations: Three implementations were compared: a fine-tuned MI2 classifier (default), a linear probe of QRad's frozen encoder, and single-token yes/no probabilities from QRad's VQA. The MI2 classifier and QRad linear probe achieve comparable overall AUC (0.90 and 0.91 respectively), while the QRad token-logit variant trails significantly (0.73).

  • Alternative LLM Rewriters: The study compared model sizes (GPT-5, GPT-5.4, GPT-5.4-mini), model providers (OpenAI and DeepSeek), and reasoning efforts. Removing per-class example text segments from the instruction degraded performance (e.g., GPT-5's AUC on Fracture dropped from.89 to.73). Reasoning effort was found to be vital, especially for smaller models; raising GPT-5.4-mini's reasoning effort to medium improved its Lung Lesion AUC from.58 to.81.

Contributions:

The paper claims three main contributions:

  1. Clinical adaptability: clinician-customizable reporting allows operating-point selection to match the clinical context.

  2. Regulatory validation: ROC-based evaluation of generated reports provides a viable path toward regulatory approval.

  3. Diagnostic accuracy: fusing complementary model types improves diagnostic accuracy over uncontrolled generation.

Extension:

The framework extends naturally to three-dimensional threshold control (Ta, Tb, Tc) per class for clinical deployment. This allows for a semi-positive zone where findings are described with explicit uncertainty and a per-class urgency score for time-sensitive findings, requiring only modified rewriting instructions rather than model retraining.

Limitations:

The paper notes that RadFusion relies on the quality of all three components: classifier errors propagate, the report generator may omit findings, and the LLM rewriter may introduce subtle linguistic artifacts. The rewriting instruction is template-based and tuned for chest X-ray findings, requiring re-tuning for other domains. Threshold control operates over a predefined set of disease classes, and findings outside this set are not subject to adjustment. The evaluation relies on automatic label extraction, which may introduce error. The system is intended to assist, not replace, radiologist review.

Improvements for AI systems

Based on this paper, here are the specific improvements I can implement in AI systems:

  • Improvement: Add a tunable sensitivity/specificity parameter (τ) to any generative AI system that produces diagnostic or classification-based text output.

  • What the improved system can do: A clinician can set τ=0.2 for emergency triage (catch all possible findings, accept false positives) or τ=0.8 for confirmatory diagnosis (only report high-confidence findings). The system generates a report whose stated diagnoses match the chosen operating point, without retraining.

  • Improvement: Implement a validation protocol that converts generated text back into structured labels and traces an ROC curve across thresholds.

  • What the improved system can do: Before deployment, the system can be validated against regulatory standards (e.g., FDA clearance) by demonstrating that its text-based diagnoses achieve a controllable, monotonic sensitivity-specificity trade-off—not just a single fixed accuracy metric.

  • Improvement: Combine a discriminative classifier (for precise per-class confidence) with a generative VQA model (for rich, open-vocabulary descriptions) and an LLM rewriter (for alignment).

  • What the improved system can do: The system leverages the classifier's numeric precision and the generator's descriptive richness. For example, if the classifier says pneumothorax: 0.85 and the generator says large left apical pneumothorax with lung collapse, the rewriter produces a report that states the finding confidently at τ=0.3 but hedges it at τ=0.7, while preserving the anatomical detail.

  • Improvement: Extend from binary thresholding to three thresholds (Ta, Tb, Tc) per class, creating a semi-positive zone.

  • What the improved system can do: For a class with confidence between Ta and Tb, the report uses language like possible early infiltrate, recommend follow-up instead of definitive statements. For time-sensitive findings (e.g., tension pneumothorax), the system can add urgency markers when confidence exceeds Tc, enabling prioritized radiology worklists.

  • Improvement: Dynamically adjust the LLM rewriter's reasoning effort based on model size and task complexity.

  • What the improved system can do: For a small model (e.g., GPT-5.4-mini), the system automatically increases reasoning effort to medium or high when handling ambiguous cases (e.g., lung lesions), improving AUC from 0.58 to 0.81. For large models, it uses lower effort for speed, saving compute without sacrificing accuracy.

  • Improvement: Embed per-class example text segments into the rewriting prompt, rather than generic instructions.

  • What the improved system can do: For fracture detection, the prompt includes an example like If classifier says fracture=0.9, report 'acute comminuted fracture of the distal radius with cortical disruption'. This raises AUC from 0.73 to 0.89, reducing linguistic ambiguity in how the LLM translates numeric scores into clinical language.

  • Improvement: Use the Auto-VQA mechanism to query the generator for missing details when the classifier detects a positive class not mentioned in the initial report.

  • What the improved system can do: If the classifier flags cardiomegaly but the generator's first pass omitted it, the system automatically issues a follow-up query (Describe the cardiac silhouette size and contour) and incorporates the response, ensuring no clinically significant finding is dropped.

  • Improvement: Replace the chest X-ray-specific rewriting template with a parameterized, domain-agnostic template.

  • What the improved system can do: The system can be re-tuned for other domains (e.g., pathology slides, retinal scans) by swapping the class list and example segments, without redesigning the fusion architecture. This reduces deployment time from months to days for new clinical specialties.

Abstract

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report can neither adapt to these scenarios nor support the ROC-based validation widely expected for regulatory clearance. We introduce RadFusion, a framework that equips report generation with threshold controllability. Our method fuses a multi-label classifier, which provides per-disease confidence scores, with a VQA-based report generator, which describes medical findings in detail; an LLM then rewrites the report so that its stated diagnoses follow the classifier's decisions at the selected threshold while staying grounded in the generator's descriptions. On MIMIC-CXR, the performance of RadFusion conforms to the classifier's ROC curve: sweeping the threshold and mapping the reports back to class labels reproduces the classifier's validated ROC performance. This conformance makes generated reports quantitatively evaluable through ROC analysis, strengthening the case for regulatory clearance, and enables operating-point selection that matches report behavior to clinical context. Moreover, combining the two model types improves diagnostic accuracy over uncontrolled generation: sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. These results show that RadFusion makes report generation clinically adaptable, quantitatively verifiable, and diagnostically more reliable.

Sources

Related papers