Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

arXiv:2608.12590 · cs.AI, cs.CV · Submitted 2026-08-12 · Read on arXiv

Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li

Sun Yat-sen University · Zhejiang University · University of Pennsylvania · Southern Medical University · Zhejiang Normal University

cs.AI, cs.CV

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: Under review

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 95/100

The gist: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support

Terminology

Summary

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21% for nodule segmentation and a mean AUROC of 0.9466 for benign malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3% to 86.2%, and reduced segmentation and reporting time by 35.9% and 27.4%, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

Improvements for AI systems

Improvements to AI systems:

  1. Integrate multi-task coordination into a single agentic pipeline – Instead of separate models for segmentation, classification, and reporting, build an agent that orchestrates all tasks sequentially, storing outputs in a structured, auditable evidence record per case.

  2. Add clinician-interactive correction loops – Design the AI to accept real-time human feedback on intermediate outputs (e.g., nodule boundaries or risk scores) and propagate corrections to downstream tasks (e.g., report generation), enabling a human-in-the-loop workflow.

  3. Implement evidence-grounded report assembly – Replace end-to-end language-model generation with a template-based assembly that pulls from verified diagnostic outputs (e.g., segmentation metrics, classification probabilities) to reduce hallucination and improve clinical consistency.

  4. Introduce a lesion-level semantic metric (ThyClinScore) – Use this metric during training and validation to optimize for clinically meaningful semantic accuracy, not just pixel or label overlap, aligning model performance with radiologist judgment.

  5. Enable cross-task transfer learning – Leverage shared representations from the multicentre dataset (0.3M images, 24k reports) to jointly train auxiliary tasks (lymph-node metastasis, follicular vs. papillary classification) with the primary nodule tasks, improving generalization across rare subtypes.

  6. Add location-aware language-model judge – Use a judge model that evaluates generated reports against ground truth with spatial awareness (e.g., nodule position in ultrasound), enabling automated, scalable quality control during deployment.

What the improved AI system can do:

  • Automatically segment thyroid nodules, classify benign/malignant status, predict lymph-node metastasis, and distinguish follicular vs. papillary carcinoma in a single pass, with all outputs logged for audit.

  • Allow a radiologist to correct any intermediate step (e.g., adjust a segmentation boundary) and instantly regenerate the risk score and final report, preserving consistency.

  • Generate structured, evidence-based reports that cite specific image regions and diagnostic metrics, reducing report errors and improving inter-reader agreement (from 70.3% to 86.2% consistency).

  • Reduce clinician workload by cutting segmentation time by 36% and reporting time by 27%, while maintaining high accuracy (Dice 87.2%, AUROC 0.947 for classification).

  • Self-improve over time by using ThyClinScore and the location-aware judge to flag low-confidence cases for human review, then fine-tune on corrected examples.

Abstract

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

Sources

Related papers