CARE: Confidence-Aware Reasoning for Reliable Medical VQA

arXiv:2608.10964 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu

Zhejiang University · Ant Group · University of Michigan · City University of Hong Kong

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted by MICCAI 2026

Code: https://github.com/anotherbricki/CARE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: CARE: Confidence-Aware Reasoning for Reliable Medical VQA proposes a framework to address confidence miscalibration in medical Multimodal Large Language Models (MLLMs) used for visual question

Terminology

Summary

CARE: Confidence-Aware Reasoning for Reliable Medical VQA proposes a framework to address confidence miscalibration in medical Multimodal Large Language Models (MLLMs) used for visual question answering (VQA). The paper states that existing models suffer from confidence miscalibration—a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. CARE jointly optimizes accuracy and calibration through a dual-stage pipeline: first, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning; second, Group Relative Policy Optimization (GRPO) with a novel Confidence-Aware Reward (CAR) mechanism ties the model’s confidence to diagnostic correctness within the reward signal. The main contributions are: (1) Confidence-Aware Reinforcement Fine-Tuning, integrating CAR into GRPO to align expressed confidence with diagnostic accuracy; (2) Scalable Medical-CoT Data Construction, an automated synthesis pipeline for high-quality medical reasoning data with verifiable conclusions; (3) Consistent Improvements in Accuracy and Calibration, with evaluations across multiple benchmarks showing superior accuracy and reduced miscalibration. The method involves a reverse-thinking synthesis pipeline where a base MLLM generates reasoning trajectories conditioned on ground truth, filtered by a verifier for logical consistency. Training proceeds in two phases: SFT cold start on the curated CoT data, followed by GRPO-based RL. The CAR reward comprises format, output, and calibration components, where calibration reward is computed as Rcalib = Rout · C(ai) − λ(1 − Rout) · C(ai), with C(ai) being the average token probability over the answer span, rewarding confident correct predictions and penalizing overconfident incorrect ones. Experiments on VQA-RAD, SLAKE, and PathVQA show CARE achieves the highest accuracy (e.g., 0.873 on SLAKE, 0.767 on VQA-RAD) while obtaining the lowest Expected Calibration Error (e.g., 0.115 on SLAKE, a 36% relative improvement over the second-best) and lowest Hallucination Rate (e.g., 0.048 on VQA-RAD). Ablation studies reveal that for closed-ended questions, RL alone performs best (e.g., 0.864 ACC and 0.096 ECE on VQA-RAD), while for open-ended questions, SFT+RL dominates (e.g., 0.881 on SLAKE Open), with SFT providing the structured CoT foundation and RL refining calibration. The paper concludes that CARE achieves the best diagnostic accuracy while simultaneously obtaining the lowest ECE and Hallucination Rate, demonstrating that accuracy and calibration can be jointly improved rather than traded off.

Improvements for AI systems

Improvements to AI systems:

  1. Confidence-Aware Reward Integration in RLHF/RLVR Pipelines
  • Modify the reward function in reinforcement learning (e.g., GRPO, PPO) to include a calibration term that penalizes overconfident incorrect answers and rewards confident correct ones, using token-level probability as a proxy for confidence. This directly reduces miscalibration during training, not just at inference.
  1. Reverse-Thinking CoT Data Synthesis for Verifiable Reasoning
  • Implement a pipeline where a base model generates reasoning trajectories conditioned on the ground truth (reverse thinking), then filters them with a verifier for logical consistency. This produces high-quality, structured chain-of-thought data without manual annotation, enabling scalable cold-start fine-tuning for domains with scarce expert labels.
  1. Dual-Stage Training (SFT + Confidence-Aware RL) with Question-Type Adaptation
  • Use a two-phase training: first SFT on the synthesized CoT data, then RL with the confidence-aware reward. Dynamically select the training regime based on question type—RL-only for closed-ended (e.g., binary) questions, SFT+RL for open-ended ones—to optimize both accuracy and calibration per task.
  1. Hallucination Reduction via Calibration-Driven Reward Shaping
  • Leverage the calibration reward component (Rcalib) to explicitly suppress hallucinated answers (high-confidence but wrong outputs). This provides a trainable mechanism to reduce hallucination rates in medical VQA and can be generalized to other high-stakes domains (e.g., legal, finance) where false confidence is dangerous.
  1. Unified Metric Optimization for Accuracy and Calibration
  • Replace single-objective training (accuracy only) with a joint objective that combines accuracy and Expected Calibration Error (ECE) or Brier score. This ensures the model does not trade off one for the other, as demonstrated by CARE’s simultaneous improvements.

What the improved AI system can do:

  • Provide reliable confidence scores in medical visual question answering, enabling clinicians to trust or reject AI suggestions based on calibrated certainty (e.g., flag low-confidence cases for human review).

  • Generate structured, verifiable reasoning paths for medical images without costly expert annotations, accelerating deployment in new specialties or rare diseases.

  • Adapt its training strategy per question type—e.g., being conservative for yes/no questions (RL-only) and more exploratory for open-ended ones (SFT+RL)—to maximize both correctness and calibration.

  • Reduce hallucinations in high-stakes outputs, such as misdiagnosing a condition with high confidence, by penalizing overconfident errors during training.

  • Serve as a general framework for any multimodal QA system (e.g., radiology, pathology, dermatology) where both accuracy and trustworthy uncertainty are critical, with measurable improvements in ECE and hallucination rates over existing MLLMs.

Sources

Related papers