How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
Yiheng Xiong, Luisa Gallée, Daniel Santak Wolf, Heiko Hillenhagen, Michael Götz
Ulm University Medical Center · Ulm University
cs.CV, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This paper evaluates the complete unsupervised domain adaptation (UDA) pipeline in medical imaging, considering both the adaptation stage and the label-free model selection stage together, under
Terminology
Summary
This paper evaluates the complete unsupervised domain adaptation (UDA) pipeline in medical imaging, considering both the adaptation stage and the label-free model selection stage together, under clinical deployment conditions. The authors argue that deploying UDA in clinical practice requires a practitioner to decide which algorithm to use and which of its checkpoints to ship, yet because the deployment (target) domain is unlabeled, models cannot be evaluated directly on it, leaving it unclear which to select.
The study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets across brain MRI, chest X-ray (CXR), and retinal imaging. For brain MRI, five UDA scenarios are constructed using ADNI-1, ADNI-2, ADNI-3, and AIBL datasets (ADNI-1→ADNI-2, ADNI-1→ADNI-3, ADNI-2→ADNI-1, ADNI-2→ADNI-3, and ADNI-1+2→AIBL). For CXR, two transfers are used in both directions (RSNA↔Child CXR and LDD↔CRD) using four datasets. For retinal imaging, cross-modality transfers between SLO and OCT are evaluated using the FairDomain dataset. The tasks are binary classification: Alzheimer's disease vs. cognitively normal for brain MRI, pneumonia vs. non-pneumonia for CXR, and glaucoma vs. non-glaucoma for retinal imaging.
The evaluation includes ten UDA algorithms spanning multiple paradigms: feature-distance minimization (MMD), adversarial alignment (DANN, CDAN, DALN), information maximization (MCC), SVD loss (BNM), pseudo-labeling (ATDOC), classifier discrepancy (MCD), and medical-specific techniques (AD2A for brain MRI, CoUDA for CXR). It also includes 13 label-free selection methods (validators): source-guided criteria (Source-Risk, IWCV, DEV, DEV-N) and target-based ones (Entropy, InfoMax, Corr-C, BNM (V), MCC (V), SND, ClassAMI, MixVal, TransScore). Each validator assigns a scalar validation score to each checkpoint without using target labels, and the checkpoint with the best score is selected. In total, this amounts to roughly 16,500 checkpoint configurations (each set by the algorithm, its hyperparameters, and training iteration), or over 80,000 trained checkpoints once repeated across folds or random seeds.
The experimental setup uses a three-layer MLP classification head with dropout rate 0.5. For brain MRI, a 3D ResNet-50 trained from scratch is used as the backbone with batch size eight per domain. For CXR and retinal imaging, a 2D ResNet-50 pretrained on ImageNet is used with batch size 48 per domain. All algorithms are trained with AdamW (weight decay 1e-4) and a one-cycle learning rate schedule with warm-up and a peak learning rate of 1e-3, for 10k iterations on brain MRI and retinal imaging and 30k iterations on CXR. The adaptation strength λ is varied over 0.1, 0.5, 1.0, giving three runs per algorithm. After warm-up, checkpoints are saved at uniform intervals, yielding 50 checkpoints per run and 150 checkpoints per algorithm. Target performance is measured on the target validation set using balanced accuracy.
The paper's main findings are as follows:
Adaptation works in principle. Under Oracle selection (which uses target labels), the Across-Algo selection (pooling checkpoints from all algorithms) exceeds SourceOnly in every scenario and approaches TargetOnly. Averaged over all scenarios, the ordering is consistent: SourceOnly < Avg. < Across-Algo, with Across-Algo reaching 84.8% and approaching TargetOnly (87.5%). A capable adapted model usually exists among the candidates.
There is a large selection gap from validators. The models selected by the evaluated validators leave a large target performance gap to the best available model. In every scenario, both Best Validator and Best Pair fall below the Oracle: the gap reaches up to 10.5 points (RSNA→Child CXR) and averages 6.1 points across scenarios. Crucially, Best Validator and Best Pair differ from scenario to scenario and are themselves identified using target labels; at deployment, where no labels are available to choose them in advance, the realized gap can only be larger.
The selection gap persists across architectures. The authors repeat RSNA→Child CXR with four backbones (ResNet-50, ResMLP, ConvNeXt, DeiT) and find the gap persists across all four, suggesting it is not specific to a particular backbone.
The selection gap is structural. The authors trace the origin of the gap to validator reliability, measured by the Spearman rank correlation ρ between validation scores and true target accuracy. For within-algorithm selection, a suitable validator may exist among evaluated ones, but which one differs by algorithm and scenario, so it cannot be chosen in advance. For example, BNM (V) ranks DALN's checkpoints well (ρ = 0.88) yet ranks ATDOC's checkpoints in reverse (ρ = −0.42) on the same scenario. Similarly, ClassAMI correlates positively for ADNI-2→ADNI-1 (ρ = 0.59) but reversed for SLO→OCT (ρ = −0.51). For across-algorithm selection, the correlation collapses to near zero (ρ = −0.04 for Source-Risk on ADNI-1→ADNI-2), and most validators are unreliable or contradict their intended direction, with correlations weak or negative. The gap therefore appears largely structural rather than a matter of choosing a better validator among those evaluated.
Two strategies narrow the gap but do not close it. The first strategy ensembles predictions across validator-selected checkpoints from every algorithm. Ensembling matches or exceeds Best Val. in most scenarios, with the remaining gap to the Oracle small in several cases (e.g., 0.6 for OCT→SLO and 2.6 for LDD→CRD), but it remains far from the Oracle in other cases such as RSNA→Child CXR. The second strategy spends a small target-labeling budget (5 to 25 labeled samples for brain MRI, 1% to 5% for other scenarios). A small budget improves selection steadily, surpassing both Best Val. (78.7) and ensembling (78.4) from B4 (20 labeled samples for brain MRI, 4% for the other scenarios) on average, though a gap to the Oracle still remains.
The paper also notes that validator reliability, measured by Spearman correlation ρ, and the target accuracy of the selected model do not always align. For example, for ADNI-1→ADNI-2, ClassAMI has a higher correlation than InfoMax (ρ = 0.24 vs. 0.16) yet selects a less accurate model (80.8% vs. 83.2%). This is because ρ captures how well a validator orders all checkpoints, whereas selection depends on its single top-scored checkpoint.
The authors conclude that the main barrier between current UDA and clinical deployment lies less in producing a capable adapted model, which usually exists, than in selecting it without target labels. They state: Deployment therefore depends on the complete pipeline, not adaptation alone: selection remains a largely unsolved step, and closing it would bring much of current UDA closer to clinical use.
They also note that in some scenarios the achievable accuracy itself remains low, and for some algorithms even the Oracle-selected model falls below SourceOnly, so better adaptation algorithms are still needed.
The paper's limitations include that it covers only binary classification; extending to multi-class, multi-label, segmentation, or detection tasks is a valuable future direction. It also evaluates by balanced accuracy, while other metrics such as sensitivity and specificity also matter clinically.
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:
1. Add a selection-aware
training objective that optimizes for validator reliability, not just target accuracy.
What the improved system can do: During UDA training, instead of only minimizing task loss and domain alignment loss, it can also maximize the Spearman correlation between a chosen validator’s scores and the model’s true target accuracy on a held-out pseudo-target set. This directly addresses the structural selection gap by making the model’s checkpoints more rankable by label-free validators, reducing the risk of a validator ranking the best checkpoint last.
2. Implement a meta-validator that dynamically selects the best validator per algorithm and scenario using only source-domain metadata.
What the improved system can do: Since no single validator works across all algorithms (e.g., BNM works for DALN but reverses for ATDOC), the system can learn a meta-policy that, given the algorithm type, hyperparameters, and training dynamics (e.g., loss curves, gradient norms), predicts which validator will have the highest ρ for that specific run. This meta-validator is trained on source-domain simulations (where target labels are available for validation) and then applied at deployment, narrowing the gap without any target labels.
3. Introduce an ensemble-of-validators
selection strategy that combines multiple validators’ top-k checkpoints into a single robust choice.
What the improved system can do: Instead of trusting one validator, the system can take the union of the top-3 checkpoints from each of the 13 validators, then use a consensus voting or rank-aggregation method (e.g., Borda count) to pick the final checkpoint. This mimics the paper’s finding that ensembling predictions across validator-selected checkpoints helps, but it does so at the checkpoint-selection stage, reducing variance and avoiding the worst-case single-validator failure.
4. Add a small-budget active selection
module that optimally spends a tiny label budget (e.g., 5–25 samples) to correct validator bias.
What the improved system can do: The system can first rank checkpoints using all validators, then actively query target labels for the most uncertain or most divergent checkpoints (e.g., those where validators disagree most). With just 1–5% labeled samples, it can re-rank the top candidates and close most of the gap to the Oracle, as the paper shows this budget is highly effective. This makes deployment practical in clinical settings where a few labels are feasible.
5. Build a deployment-risk estimator
that predicts when the selection gap will be large (e.g., >5 points) using only source-domain and validator statistics.
What the improved system can do: Before deployment, the system can compute features like (a) the variance of validator scores across checkpoints, (b) the pairwise disagreement among validators, and (c) the source-domain accuracy of the best checkpoint. If these features indicate high selection risk, the system can flag the scenario as not deployment-ready
and recommend either collecting a small label budget or using a more conservative ensemble approach, rather than shipping a potentially poor model.
6. Extend the selection pipeline to multi-class and segmentation tasks by using task-agnostic validators (e.g., confidence calibration, feature-space density) that do not rely on binary classification assumptions.
What the improved system can do: The current validators are designed for binary classification. The improved system can use generalizable validators like (a) the negative entropy of the softmax distribution (works for any number of classes), (b) the distance-to-training-distribution in feature space (works for segmentation), and (c) the consistency of predictions under input perturbations (e.g., test-time augmentation). This allows the same selection framework to be applied to multi-class disease classification, lesion segmentation, or detection tasks, which are common in clinical practice.
7. Create a validator-aware checkpoint scheduler
that saves checkpoints at training iterations where validator scores are most informative.
What the improved system can do: Instead of saving checkpoints at uniform intervals, the system can monitor validator scores during training and save checkpoints at moments when validators show high disagreement or when the top-ranked checkpoint changes rapidly. This yields a more diverse and informative candidate pool, increasing the chance that a good model exists among the saved checkpoints, and reducing the risk that the best model is missed between saves.
Summary of what the improved AI system can do:
It can deploy UDA models in clinical settings with a much smaller performance gap to the oracle, by (a) making models more rankable by validators, (b) intelligently combining validators, (c) using a tiny label budget effectively, (d) flagging risky deployment scenarios, and (e) extending selection to more complex medical tasks—all without requiring target labels at scale.
Sources
- SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities
- Medical Image Segmentation with Domain Adaptation: A Survey
- Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation
- Unsupervised Domain Adaptation: A Reality Check
- Three New Validators and a Large-Scale Benchmark Ranking for Unsupervised Domain Adaptation
- Domain Adaptation and Generalization on Functional Medical Images: A Systematic Survey
- M3DA: Benchmark for Unsupervised Domain Adaptation in 3D Medical Image Segmentation
- Navigating Distribution Shifts in Medical Image Analysis: A Survey
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models