MIRA: Medical Image Reflection for Agentic Diagnosis

arXiv:2608.10827 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu

Tongji University · Nanjing University · Northwestern Polytechnical University · People's Public Security University of China · École Normale Supérieure – PSL

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Project page: https://mira-vl.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: MIRA (Medical Image Reflection for Agentic Diagnosis) is a medical visual diagnostic framework that enhances Large Vision-Language Models (LVLMs) with active evidence acquisition, tool-grounded

Terminology

Summary

MIRA (Medical Image Reflection for Agentic Diagnosis) is a medical visual diagnostic framework that enhances Large Vision-Language Models (LVLMs) with active evidence acquisition, tool-grounded verification, and reflection-driven self-correction. The paper addresses the challenge that indiscriminate tool use may introduce noisy observations, propagate misleading evidence, and increase reasoning cost, making it crucial for medical agents to decide when external tools are necessary. The framework is guided by four core principles: evidence-grounded visual interaction, tool-use reliability, failure-aware reflection, and continual principle evolution.

The technical approach comprises a two-stage training strategy. In the first stage, Supervised Fine-Tuning (SFT) is conducted on a curated medical tool-use and reflection dataset. This dataset is built from the training splits of four public medical VQA datasets (PMC-VQA, SLAKE, VQA-RAD, PathVQA) plus an in-house dataset. The SFT stage uses a tool-augmented Monte Carlo Tree Search (MCTS) data engine that systematically explores the reasoning space by generating diverse diagnostic hypotheses, and performs joint verification of visual grounding accuracy and semantic reasoning consistency at each node during tree expansion. The MCTS is configured with 20 simulations, maximum depth of 4, branching factor of 3, and exploration constant cpuct = 1.4. The SFT also includes reflection data construction, which synthesizes both failure-driven correction data (from wrong trajectories) and textual intention data (from successful trajectories).

In the second stage, Reinforcement Learning (RL) is applied using GRPO with a composite trajectory reward. The reward includes three components: formatting reward (Rformat), result reward (Rresult), and consistency reward (Rcons). The consistency reward evaluates whether the final answer is supported by the preceding evidence and reasoning. The RL stage also introduces an online reflective principle evolution mechanism where failure cases are distilled into reflective principles and injected into subsequent rollouts, with effective principles selected through a trial-and-verification loop. The reflection memory is updated only when it improves held-out validation reward, using a validation-gated mechanism.

The model uses six tools: SEARCH (web knowledge retrieval), GROUNDING (bounding-box localization), POINT (point-rendering for fine-grained evidence), ZOOM (fine-grained inspection), ROTATE (orientation adjustment), and MEASURE (relative measurement). All tools use a normalized 0-999 coordinate convention.

Experimental results show that MIRA-VL-8B improves the average score from 57.29 to 64.73 compared to the Qwen3-VL-8B backbone across nine medical visual reasoning benchmarks. The gains are especially pronounced on MedLesionVQA and MedLesionMCQ, with improvements of 14.5 and 14.4 points respectively. Fine-grained analysis shows that improvements mainly come from Basic Perception and Understanding/Diagnosis/Suggestion categories, while Content Recognition changes only marginally. The tool-use necessity analysis shows that MIRA-VL-8B increases useful tool use from 56.2% to 73.8% and reduces harmful tool use from 8.9% to 1.6% compared to direct tool access for the backbone.

The paper acknowledges limitations including base model capability constraints and reflection memory generalization challenges. The authors note that MIRA is still constrained by the perception, language understanding, and medical knowledge of its backbone model and that the learned memories may not fully generalize to rare diseases, unseen imaging styles, or substantially different clinical contexts. The SFT stage requires approximately 113 GPU hours on 40 A800 GPUs, while the RL stage requires approximately 1200 GPU hours on 32 NVIDIA H20 GPUs.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Tool-Use Gating with Cost-Aware Decisioning
  • Implement a learned “tool necessity estimator” that predicts whether external tool invocation will improve diagnostic confidence before execution, based on current visual evidence and reasoning state.

  • Integrate a dynamic budget controller that allocates tool calls based on task difficulty, penalizing unnecessary calls during inference to reduce latency and computational cost.

  • Improved system capability: Medical AI can autonomously decide when to query web knowledge, zoom into regions, or measure structures, avoiding noisy or misleading tool outputs while maintaining high diagnostic accuracy under real-time clinical constraints.

  1. Reflective Principle Memory with Cross-Domain Generalization
  • Extend the reflection memory to a hierarchical, graph-structured store where principles are indexed by disease category, imaging modality, and anatomical region, enabling retrieval of relevant past failures for unseen cases.

  • Add a meta-learning loop that periodically re-evaluates stored principles against new data distributions, pruning outdated or overfit rules and synthesizing new ones via contrastive analysis of failure clusters.

  • Improved system capability: The AI can transfer learned correction strategies from common conditions (e.g., pneumonia) to rare or novel presentations (e.g., atypical fungal infections) by retrieving analogous structural or semantic patterns, reducing performance degradation on out-of-distribution imaging.

  1. Joint Visual-Semantic Consistency Verification via Multi-Agent Debate
  • Replace single-path consistency reward with a dual-agent architecture where one agent generates diagnostic hypotheses and another independently verifies evidence alignment, using adversarial feedback to refine reasoning.

  • Incorporate a “counterfactual probe” that perturbs visual grounding (e.g., shifting bounding boxes or zoom regions) and measures answer stability, flagging cases where conclusions hinge on fragile evidence.

  • Improved system capability: The AI can self-audit its reasoning chains, rejecting conclusions that are not robustly supported by visual details, thereby reducing false positives in lesion detection and improving trustworthiness in high-stakes diagnostic settings.

  1. Continual Principle Evolution with Validation-Gated Memory Updates
  • Deploy an online learning pipeline where the reflection memory is updated during deployment using a shadow validation set from each new clinical site, with updates only committed if they improve site-specific reward metrics.

  • Add a “principle conflict resolver” that detects when newly distilled principles contradict existing ones, triggering a re-simulation of conflicting cases via MCTS to resolve ambiguity before memory integration.

  • Improved system capability: The AI continuously adapts to new imaging protocols, scanner types, and population demographics without catastrophic forgetting, maintaining high accuracy across heterogeneous clinical environments.

  1. Tool-Grounded Failure Attribution for Rare Disease Handling
  • Introduce a failure attribution module that, when the final answer is incorrect, traces back to identify whether the error originated from visual grounding, tool output, or reasoning, and generates targeted corrective data for each failure type.

  • Use this attribution to augment the SFT dataset with synthetic rare-disease cases generated by perturbing existing tool outputs (e.g., simulating atypical lesion shapes via MEASURE or ZOOM variations) and applying reflection principles.

  • Improved system capability: The AI can better handle rare diseases by learning from decomposed failure modes, improving diagnostic accuracy for conditions with limited training data, and providing explainable error analysis for clinicians.

  1. Efficient Multi-Tool Coordination with Hierarchical Planning
  • Implement a two-level planner: a high-level policy selects which tool category (e.g., SEARCH vs. GROUNDING) to use based on the current reasoning stage, and a low-level policy decides specific tool parameters (e.g., zoom factor, rotation angle) via a lightweight neural network trained with the MCTS data.

  • Add a “tool fusion” mechanism that combines outputs from multiple tools (e.g., GROUNDING + MEASURE) into a unified evidence tensor, reducing redundant calls and improving feature richness for downstream reasoning.

  • Improved system capability: The AI can perform complex diagnostic workflows (e.g., locate a lesion, measure its size, search for matching clinical knowledge) in fewer steps, reducing inference time by 30% while maintaining or improving accuracy on multi-step visual QA tasks.

Abstract

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/

Sources

Related papers