Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
Sun Yat-sen University · Zhejiang University · University of Pennsylvania · Southern Medical University · Zhejiang Normal University
cs.AI, cs.CV
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: Under review
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 95/100
The gist: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support
Terminology
Summary
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21% for nodule segmentation and a mean AUROC of 0.9466 for benign malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3% to 86.2%, and reduced segmentation and reporting time by 35.9% and 27.4%, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
Improvements for AI systems
Improvements to AI systems:
-
Integrate multi-task coordination into a single agentic pipeline – Instead of separate models for segmentation, classification, and reporting, build an agent that orchestrates all tasks sequentially, storing outputs in a structured, auditable evidence record per case.
-
Add clinician-interactive correction loops – Design the AI to accept real-time human feedback on intermediate outputs (e.g., nodule boundaries or risk scores) and propagate corrections to downstream tasks (e.g., report generation), enabling a
human-in-the-loop
workflow. -
Implement evidence-grounded report assembly – Replace end-to-end language-model generation with a template-based assembly that pulls from verified diagnostic outputs (e.g., segmentation metrics, classification probabilities) to reduce hallucination and improve clinical consistency.
-
Introduce a lesion-level semantic metric (ThyClinScore) – Use this metric during training and validation to optimize for clinically meaningful semantic accuracy, not just pixel or label overlap, aligning model performance with radiologist judgment.
-
Enable cross-task transfer learning – Leverage shared representations from the multicentre dataset (0.3M images, 24k reports) to jointly train auxiliary tasks (lymph-node metastasis, follicular vs. papillary classification) with the primary nodule tasks, improving generalization across rare subtypes.
-
Add location-aware language-model judge – Use a judge model that evaluates generated reports against ground truth with spatial awareness (e.g., nodule position in ultrasound), enabling automated, scalable quality control during deployment.
What the improved AI system can do:
-
Automatically segment thyroid nodules, classify benign/malignant status, predict lymph-node metastasis, and distinguish follicular vs. papillary carcinoma in a single pass, with all outputs logged for audit.
-
Allow a radiologist to correct any intermediate step (e.g., adjust a segmentation boundary) and instantly regenerate the risk score and final report, preserving consistency.
-
Generate structured, evidence-based reports that cite specific image regions and diagnostic metrics, reducing report errors and improving inter-reader agreement (from 70.3% to 86.2% consistency).
-
Reduce clinician workload by cutting segmentation time by 36% and reporting time by 27%, while maintaining high accuracy (Dice 87.2%, AUROC 0.947 for classification).
-
Self-improve over time by using ThyClinScore and the location-aware judge to flag low-confidence cases for human review, then fine-tune on corrected examples.
Abstract
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
Sources
- CLIP-TNseg: A Multi-Modal Hybrid Framework for Thyroid Nodule Segmentation in Ultrasound Images
- Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge
- MedSAM2: Segment Anything in 3D Medical Images and Videos
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MedGemma Technical Report
- DINOv3
- Long-tail learning via logit adjustment
- AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data
- Qwen3-VL Technical Report
- GPT-4o System Card
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection