Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases

arXiv:2602.03302 · cs.CV, cs.AI · Submitted 2026-02-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases".

Jane: The paper was written by Jinze Zhang, Jian Zhong, Li Lin, Jiaxiong Li, Ke Ma et al. from State Key Laboratory of Ophthalmology, Zhongshan Ophthalmic Center, Sun Yat-sen University and Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science, Guangzhou, China and Department of Electrical and Electronic Engineering, Southern University of Science and Technology and Department of Electrical and Electronic Engineering, The University of Hong Kong (Hong Kong SAR) and Department of Health Technology and Informatics, The Hong Kong Polytechnic University and Eye Center, Zhongshan City People's Hospital and Zhongshan, Guangdong, China and Guangdong Medical University and Zhanjiang, Guangdong, China and The Third Affiliated Hospital of Sun Yat-sen University and Guangzhou, China and Department of Ophthalmology, Zhaoqing Gaoyao People’s Hospital and Zhaoqing, China and Tianjin Key Laboratory of Retinal Functions and Diseases, Tianjin Branch of National Clinical Research Center for Ocular Disease, Eye Institute and School of Optometry, Tianjin Medical University Eye Hospital and Tianjin, China and Jiaxing Research Institute, Southern University of Science and Technology and Jiaxing, China and Beijing Tongren Eye Center, Beijing Tongren Hospital, Capital Medical University and Beijing, China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We’ve covered the scope and ambition of this work, so let's look deeper into what the paper promises in its abstract.

Jane: The summary tells us that while OCT has revolutionized diagnosis, its manual workflow has remained a constraint for large-scale screening.

Lu: This system, which they call FOCUS, is designed to directly address that limitation by creating an end-to-end solution to the entire three dee OCT process.

Meng: It’s interesting how they frame it—it's not just a diagnostic tool; it's a full-process clinical utility system that aims to automate the workflow itself.

Lalam: The promise here is that we are moving away from fragmented AI and toward a cultural shift where consistent, automated expert-level care becomes the global standard.

Tom: They detail how they achieve this automation by sequentially performing image quality assessment with EfficientNetV2-S before moving on to abnormality detection.

Jane: That initial check for image quality is crucial because it means the subsequent diagnoses are built on solid data, preventing unreliable inputs from ruining the whole process.

Lu: Then, they use a fine-tuned Vision Foundation Model for the multi-disease classification, which allows them to leverage that massive pre-trained knowledge base.

Meng: I'm interested in how they handle the transition between these stages; it sounds like a sophisticated handover process from quality control to the actual predictive models.

Lalam: This is about ensuring that if we scale this up, we aren't sacrificing diagnostic accuracy for speed, which is a massive win for healthcare delivery.

The Core Methodology: Tom: We've looked at the overall promise of FOCUS; now let's dig into the core mechanics of how they actually build this system.

Jane: The key innovation is that instead of just relying on individual slices, they use a unified adaptive aggregation method to pull predictions from 2D slices level and integrate them into a comprehensive three dee patient-level diagnosis.

Meng: That sounds like Multiple Instance Learning, which is exactly what we need to handle the complexity of combining those individual B-scans into a holistic view.

Lu: The integration of the foundation model with this workflow is what allows for such seamless and logical transitions between diagnostic stages for the AI.

Tom: It’s taking that powerful, generalizable AI model and adapting it specifically to handle the complexity of assembling those individual slices into a coherent, volumetric picture for us.

Jane: This mechanism allows us to capture those subtle, sparse features—the 'needle-in-a-haystack' signs—that might be completely missed if we only looked at each image independently.

Lalam: The cultural impact of this is that it enables a standardized approach, suggesting an automated system could provide consistent, expert quality of care regardless of where the patient is or who performs the initial screening.

Meng: This standardization is vital for scaling up population-level retinal care because we can finally move beyond the limitations of specialist access.

Lu: I find it incredibly sophisticated that they use a Unified Adaptive Aggregation Classifier, or UAAC, to dynamically weigh each slice's contribution based on its clinical diagnostic quality.

Tom: That means the system isn't just averaging predictions; it’s making smart decisions about which evidence is reliable and building the final diagnosis around those most trustworthy pieces of information.

Jane: It allows the AI to be selective, focusing only on high-confidence pathological evidence rather than getting confused by ambiguous or equivocal findings, which is a huge improvement over older attempts.

Meng: The UAAC’s ability to guide this suggests a highly scalable and deployable architecture for deployment across multiple different hardware vendors.

Lalam: This moves us decisively away from fragmented research prototypes toward a cohesive, deployable tool that speaks directly to the future of reliable healthcare delivery.

The Core Methodology: Tom: So, we’ve seen how they built this system; now let's look at the performance data in "Full end-to-end diagnostic workflow automation of three dee OCT via foundation model-driven AI for retinal diseases."

Jane: The internal dataset shows incredibly strong F1 scores—ninety-nine point zero one percent for quality assessment and ninety-seven point four six percent for abnormal detection, which is a very robust benchmark.

Meng: Beyond the internal success, the external validation is perhaps even more critical; they tested this on one thousand three hundred forty-five patients across four different-tier centers and diverse OCT devices to see if it holds up.

Lu: The system’s performance held up remarkably well under these real-world conditions where most other AI models tend to fail, which is a major indicator of its inherent robustness.

Tom: Their real-world performance was very stable, showing F1 scores consistently between ninety and ninety-five percent, which is a huge indicator of reliability for us when dealing with messy data.

Jane: It’s truly encouraging to see that the results aren't just theoretical; they are consistent when applied across different clinical environments and various devices, which proves generalizability.

Meng: Consistency is the absolute gold standard for deployment, Jane; if it performs well in diverse clinical settings, that suggests a lot about its practical viability for mass screening.

Lu: I’m particularly excited about the potential for bringing consistency to an area where human variability—especially when only looking at OCT B-scans—is notoriously high.

Tom: The real-world performance showed it matching or even exceeding expert performance in tasks like multi-disease diagnosis, which is a very strong validation of its capability.

Jane: It’s also important to note that the results show the AI can handle complex, multi-disease diagnoses effectively at the patient level, connecting all those slice findings into a comprehensive picture.

Lalam: This proves that when we move beyond fragmented models, we are moving toward a system that provides consistent, expert-level care to every single person who needs it.

Performance and Validation: Tom: We've seen the internal and external validation; now let's look at what these specific results mean for our listeners in "Full end-to-end diagnostic workflow automation of three dee OCT via foundation model-driven AI for retinal diseases."

Jane: The performance is stellar, but we also need to acknowledge that the external validation was conducted on a diverse group of one thousand three hundred forty-five patients across four different types of medical centers.

Meng: The results show that this system is robust enough to handle complex real-world data streams without getting overwhelmed or requiring excessive retraining for deployment in a specific location.

Lu: I think the key finding here is that this end-to-end approach successfully solves the problem of bridging 2D feature extraction with volumetric interpretation, allowing us to use all the spatial context.

Tom: That means we have a tool that’s not just looking at isolated pictures but using the entire three dee volume to make a highly informed diagnosis.

Jane: It proves that when it's used in the real world, this system is consistent, which is something we usually struggle with in clinical settings due to human factors.

Meng: That consistency confirms its viability for scaling up population-level screening efforts globally because it provides reliable performance regardless of local expertise.

Lalam: I believe this capability allows us to offer reliable, expert-level care even when the resources available to a clinic are severely limited.

Tom: The fact that it matches or exceeds expert performance in tasks like abnormality detection is a massive indicator of trust in its advanced capability.

Jane: It’s also crucial for us to emphasize that the the multi-disease diagnosis module works effectively, connecting all those slice findings into a comprehensive picture for the patient.

Implications and Limitations: Tom: We've covered the scope and performance; now let's talk about what this means for our listeners in "Full end-to-end diagnostic workflow automation of three dee OCT via foundation model-driven AI for retinal diseases."

Jane: The core message is that this work solves critical gaps in how we currently use AI in ophthalmology by creating a system that mirrors the real multi-stage, expert clinical process.

Meng: The practical implication is that this could dramatically increase our efficiency and potentially standardize diagnostics for population-level screening, which is a massive logistical improvement for healthcare delivery.

Lu: I'm particularly excited about the potential for bringing consistency to an area where human variability—especially when only looking at OCT B-scans—is quite high, allowing us to solve that uncertainty.

Lalam: I think the most profound impact is that it creates a blueprint for autonomous screening, moving us toward a future where access to expert-level diagnostics is significantly improved globally.

Tom: And while we've seen impressive results, it’s also important to acknowledge the limitations—that this study was primarily based on data from various Chinese medical centers.

Jane: It is a limitation, but the authors have taken steps to mitigate that by including diverse geographical and socioeconomic regions in their cohort to enhance generalizability.

Meng: Moving forward with prospective studies will be crucial for validating how it performs in mass screening environments where disease incidence is lower than those found in high-volume clinical settings.

Lalam: We can also see the path forward involves integrating things like large language models to support automated reporting and even more transparent decision-making processes.

The Future and Conclusion: Tom: We've covered so much ground today, from the authors to these impressive findings in "Full end-to-end diagnostic workflow automation of three dee OCT via foundation model-driven AI for retinal diseases."

Jane: The core message is that this work successfully mirrors the real multi-stage clinical process, creating a system that feels natural and trustworthy.

Meng: The practical implication remains the massive increase in efficiency, standardizing diagnostics for population-level screening efforts.

Lu: I'm particularly excited about the potential for bringing consistency to an area where human variability is high, ensuring every patient gets a consistent assessment.

Lalam: I think the most profound impact is that it creates a blueprint for autonomous screening, significantly improving global access to expert-level diagnostics.

Tom: We've seen a monumental shift from narrow, single-disease models to this comprehensive, multi-disease system that is truly remarkable in its scope.

Jane: It's certainly a very exciting time in ophthalmic AI, Tom; it paves the way for much more reliable and accessible care for everyone who needs it.

Meng: I look forward to seeing how these integration challenges are solved in a practical, real-world deployment scenario that is highly scalable.

Lu: I can only hope future iterations build on this foundation to fully realize the possibilities of truly autonomous diagnostics across all that awaits us.

Lalam: We'll be watching this progress with great interest as we see how it improves the culture of standardized care globally.

Tom: This entire paper, "Full end-to-end diagnostic workflow automation of three dee OCT via foundation model-driven AI for retinal diseases," represents a monumental shift in how we manage and screen for retinal conditions.

Jane: It's truly a powerful tool that will help us reach a new level of standardized, reliable care.

Jinze Zhang, Jian Zhong, Li Lin, Jiaxiong Li, Ke Ma, Naiyang Li, Meng Li, Yuan Pan, Zeyu Meng, Mengyun Zhou, Shang Huang, Shilong Yu, Zhengyu Duan, Sutong Li, Honghui Xia, Juping Liu, Dan Liang, Yantao Wei,, Xiaoying Tang, Jin Yuan, Peng Xiao

State Key Laboratory of Ophthalmology, Zhongshan Ophthalmic Center, Sun Yat-sen University · Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science, Guangzhou, China · Department of Electrical and Electronic Engineering, Southern University of Science and Technology · Department of Electrical and Electronic Engineering, The University of Hong Kong (Hong Kong SAR) · Department of Health Technology and Informatics, The Hong Kong Polytechnic University · Eye Center, Zhongshan City People's Hospital · Zhongshan, Guangdong, China · Guangdong Medical University · Zhanjiang, Guangdong, China · The Third Affiliated Hospital of Sun Yat-sen University · Guangzhou, China · Department of Ophthalmology, Zhaoqing Gaoyao People’s Hospital · Zhaoqing, China · Tianjin Key Laboratory of Retinal Functions and Diseases, Tianjin Branch of National Clinical Research Center for Ocular Disease, Eye Institute and School of Optometry, Tianjin Medical University Eye Hospital · Tianjin, China · Jiaxing Research Institute, Southern University of Science and Technology · Jiaxing, China · Beijing Tongren Eye Center, Beijing Tongren Hospital, Capital Medical University · Beijing, China

cs.CV, cs.AI

Submitted: 2026-02-03

Updated: 2026-08-25

Importance score: 79/100

The gist: Summary of "Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases" The study addresses the limitations in current automated diagnosis of Optical

Key concepts

3D OCT
Optical Coherence Tomography is the core diagnostic tool used for retinal diseases. The system analyzes the entire three-dimensional volume of the retina, rather than just isolated two-dimensional slices. This allows for a comprehensive, patient-level diagnosis.
Foundation Model-Driven AI
This methodology employs massive pre-trained AI knowledge bases to power the diagnostic workflow. It enables seamless transitions between different stages—such as image quality assessment and multi-disease classification—maintaining high accuracy.
Unified Adaptive Aggregation Classifier (UAAC)
The UAAC is a mechanism that dynamically weighs each slice's contribution to the final diagnosis. Rather than simply averaging predictions, it intelligently focuses on reliable, high-confidence pathological evidence for greater accuracy.

Terminology

Summary

Summary of Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases

The study addresses the limitations in current automated diagnosis of Optical Coherence Tomography (OCT) for retinal diseases, noting that while OCT is a cornerstone of modern ophthalmology, its full diagnostic automation is constrained by multi-stage workflows and conventional single-slice single-task AI models. The traditional manual diagnostic workflow—which requires meticulous image acquisition, targeted selection of pathology-relevant slices, and expert interpretation—restricts the use of OCT for large-scale screening. Existing AI tools are insufficient because they are designed for isolated tasks (e.g., quality classification or disease detection) and fail to replicate the integrated, end-to-end nature of expert reasoning.

A fundamental methodological challenge identified is that retinal pathologies are inherently three-dimensional. Current 2D approaches inherently treat slices in isolation, compromising accuracy by missing volumetric context, while volumetric 3D models are often limited by prohibitive computational costs and require massive volumetric annotations. Furthermore, existing foundation models (such as RETFound and VisionFM) possess a 2D architecture, creating a barrier to 3D diagnosis. Classical Multiple Instance Learning (MIL) approaches also struggle with the sparse, ‘needle-in-a-haystack’ pathological signals.

To bridge this gap, the researchers introduce FOCUS (Full-process OCT-based Clinical Utility System), an end-to-end AI framework engineered to automate the complete clinical OCT workflow. Unlike systems that optimize isolated tasks, FOCUS is architectured to mimic the comprehensive clinical pipeline, seamlessly integrating quality control, anomaly triage, and volumetric diagnosis into a unified workflow.

Methodology and Architecture:

FOCUS employs a workflow-adaptive architecture that leverages a Vision Foundation Model enhanced by a Prompt Decoder and a Unified Adaptive Aggregation Classifier (UAAC). The system utilizes the MIL paradigm to achieve robust synthesis: "The MIL paradigm within FOCUS enables the robust synthesis of patient-level diagnoses from 2D slice streams, effectively capturing sparse pathological features within large volumes without the computational constraints of 3D-CNN."

The UAAC is central to this process; it dynamically calibrate[s] the influence of each B-scan based on its clinical diagnostic contribution and associated prediction uncertainty. This selective aggregation mechanism ensures that the final patient-level diagnosis is derived from the most reliable slice-level cues, maintaining robustness even when definitive pathological signs are sparse.

Results and Validation:

The system was trained on 3,300 patients (40,672 slices) and tested/validated on an external cohort of 1,345 patients (18,498 slices) across four different-tier centers and diverse OCT devices. The performance metrics were:

  • Quality Assessment: F1 score of 99.01%.

  • Abnormality Detection: F1 score of 97.46%.

  • Patient-level Multi-disease Diagnosis: F1 score of 94.39%.

Real-world validation across the multiple centers showed stable performance (F1: 90.22%-95.24%). In human-machine comparisons, FOCUS matched or exceeded expert performance in the anomaly detection task (F1: 95.47% vs 90.91%) and in multi-disease diagnosis (F1: 93.49% vs 91.35%), while also demonstrating better efficiency.

Discussion and Limitations:

The study concludes that FOCUS is the first OCT-AI system to successfully translate the multi-step, expert-level diagnostic process into a deployable, end-to-end solution. The design of the human–machine comparison specifically mirrored the real-world, multi-stage ophthalmic workflow.

However, several limitations were noted:

  1. Geographic Concentration: The training and validation datasets were primarily collected from Chinese medical centers, which may limit the epidemiological generalizability of FOCUS to populations with different ethnic backgrounds or clinical workflows.

  2. Retrospective Bias: The retrospective nature of the evaluation means the reported performance reflects utility in high-prevalence clinical settings, but may not fully predict its behavior in mass screening scenarios where disease incidence is low.

Future work is required to integrate large language models (LLMs) to support interactive dialogue and automated report generation. Furthermore, the transition to a widely deployed clinical tool necessitates rigorous adherence to data security and ethical standards, requiring robust privacy-preserving infrastructures and compliance with evolving data protection regulations.

Improvements for AI systems

As a diligent AI researcher, I see several critical areas where the current Focus architecture can be enhanced to address its inherent limitations and maximize its potential for real-world clinical deployment. These improvements move beyond simply fixing bugs; they involve architectural evolution to ensure global applicability, interpretability, and seamless integration into advanced healthcare ecosystems.

The Improvement: Implement a Federated Learning (FL) Framework for the Vision Foundation Model (VFM) and the subsequent UAAC module. Instead of relying on centralized data aggregation from specific regional centers, FL allows model weights to be trained locally within diverse international clinical environments. Only encrypted weight updates are shared, not raw patient data.

What the Improved System Can Do:

  • Achieve Global Robustness: Eliminate the bias stemming from the current concentration of Chinese medical center data. The system will perform reliably across populations with different ethnic backgrounds, varied disease spectra (e.g., specific genetic predispositions), and diverse clinical workflows worldwide.

  • Adapt to Local Hardware: The FL framework can be optimized to ensure the model performs consistently across a wider range of OCT device generations and manufacturers, reducing performance drop-off in non-standard environments.

Sources

Related papers