VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation
Paul Fischer, Ece Ozkan
University of Basel
cs.CV, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 48/100
The gist: VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation Summary This paper introduces VIDS-Seg, an extension of the Variational Inference under Distribution
Terminology
Summary
VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation
Summary
This paper introduces VIDS-Seg, an extension of the Variational Inference under Distribution Shifts (VIDS) framework, designed for dense image segmentation with a focus on reliable uncertainty quantification under covariate shift. The motivation is the clinical problem of models trained on adult cohorts silently under-performing on pediatric patients, particularly infants, without any indication in the model's output that something has gone wrong.
Problem and Motivation
The paper addresses the challenge that a model that performs well on the population it was trained on may fail systematically when deployed on a different one.
The authors note that children are not a homogeneous population
and that a model trained broadly on 'pediatric' data can itself exhibit the same silent under-performance on specific pediatric subgroups, such as infants.
They emphasize that detecting such silent, systematic failures is not merely a technical curiosity but rather a prerequisite for the robust and safe deployment of machine learning in diverse clinical populations.
A critical limitation identified is that "most UQ methods... their desirable properties, calibration, coverage, and appropriate uncertainty magnitude, are typically established and evaluated on in-distribution test data. There is no general mechanism by which standard UQ methods should produce higher uncertainty for inputs that lie outside the training distribution."
Methodology
VIDS-Seg builds on the VIDS framework, which introduces an adaptive prior that conditions on test covariates
to increase uncertainty for inputs that deviate from the training distribution.
The key technical contributions are:
-
Decomposition of the network: The U-Net is split into a frozen embedding network (the full U-Net except its final layer) and a stochastic prediction head (a single 1×1 convolution). This makes
amortized variational inference over a lightweight prediction head
tractable, as the prediction head has only dθ = 2D + 2 parameters for binary segmentation. -
Spatial aggregation: Each image's dense embedding is reduced to a global descriptor via spatial average pooling, and the context set is summarized by concatenating the sample mean and standard deviation across images.
-
Segmentation-adapted energy function: The energy function averages log-probabilities over spatial locations
to ensure the energy scale is independent of image resolution.
-
Two-stage training: Stage 1 pre-trains the U-Net backbone with a hybrid cross-entropy/soft-Dice loss. Stage 2 freezes the backbone and jointly optimizes the inference network and stochastic prediction head using a cross-environment ELBO with L=40 synthetic environments, KL weight λKL=0.1, and variance-penalty weight τ=10−3.
Experimental Setup
The method is evaluated on two public datasets: EchoNet-Dynamic (adult, used for training) and EchoNet-Pediatric (pediatric, used for zero-shot evaluation). The pediatric cohort is stratified into five age groups: infants ([0,1) years, n=304), toddlers ([1,3) years, n=442), pre-schoolers ([3,6) years, n=702), school-age children ([6,13) years, n=1982), and teenagers ([13,18] years, n=2366).
Baselines include a Deep Ensemble of 10 U-Nets and PHiSeg. All models are trained solely on EchoNet-Dynamic and evaluated zero-shot on EchoNet-Pediatric.
Key Findings
-
Identifying the OOD subgroup: "All models achieve consistently high segmentation performance (DSC > 0.90) on toddlers, pre-schoolers, school-age children, and teenagers. In contrast, the infant group shows a pronounced and statistically significant performance drop to DSC ≈ 0.84 − 0.85 across all methods." This confirms infants as the OOD population.
-
Uncertainty quality: VIDS-Seg achieves the highest Normalized Cross-Correlation (NCC) between predicted entropy and actual prediction error for both non-infant and infant groups. Specifically:
-
Raw NCC: VIDS-Seg achieves 0.50 ± 0.14 (non-infant) and 0.52 ± 0.12 (infant), compared to Ensemble's 0.35 ± 0.15 and 0.39 ± 0.17, and PHiSeg's 0.14 ± 0.11 and 0.11 ± 0.09.
-
After temperature scaling: VIDS-Seg still leads with 0.61 ± 0.13 (non-infant) and 0.62 ± 0.11 (infant), versus Ensemble's 0.50 ± 0.15 and 0.54 ± 0.16, and PHiSeg's 0.35 ± 0.12 and 0.33 ± 0.12.
Statistical testing confirms significance: For the uncalibrated predictions, the difference is significant for both the infant group (p = 1.9 × 10−8, rank-biserial effect size rrb = 0.86) and the non-infant group (p = 1.1 × 10−170, rrb = 0.94).
-
Post-hoc calibration cannot substitute:
Temperature scaling improves NCC for all three methods... However, the ranking between methods is preserved after calibration.
This showsVIDS-Seg's advantage is not an artifact of better-scaled confidence that a simple post-hoc rescaling could replicate.
-
Downstream clinical impact: For ejection fraction (EF) estimation, VIDS-Seg achieves the lowest Mean Absolute Error (MAE) for infants with substantially lower variance. For detecting cardiac malfunction (EF < 50%) in infants, VIDS-Seg attains the highest AUROC of 0.94, compared to 0.90 for the Ensemble and 0.81 for PHiSeg.
Conclusion
The paper concludes that "OOD-aware uncertainty quantification can serve as a practical safety layer for deployed segmentation models, enabling detection of silent failures in underrepresented subgroups without retraining or additional labeled data. The authors state:
VIDS-Seg takes a step towards clinical ML that can communicate its own limitations. Not by eliminating distributional gaps, but by making them visible."
Improvements for AI systems
Improvements to AI systems:
-
OOD-aware uncertainty calibration: Implement VIDS-Seg's adaptive prior mechanism that conditions on test covariates, enabling the model to systematically increase uncertainty for inputs deviating from training distribution—unlike standard UQ methods that only guarantee calibration in-distribution.
-
Lightweight stochastic head decomposition: Split any dense prediction network into a frozen embedding backbone and a stochastic 1×1 convolution head (with only 2D+2 parameters for binary segmentation), making amortized variational inference tractable without retraining the full model.
-
Resolution-invariant energy function: Use spatial-average-pooled log-probabilities as the energy term, ensuring uncertainty magnitude remains comparable across images of varying resolutions—critical for multi-site or multi-device deployment.
-
Cross-environment ELBO training: Train with L=40 synthetic environments and a variance-penalty term (τ=10−3) to explicitly penalize overconfidence on shifted inputs, producing uncertainty estimates that correlate with actual error (NCC 0.61 vs 0.50 for ensembles after calibration).
-
Two-stage training protocol: Pre-train the backbone with hybrid cross-entropy/soft-Dice loss, then freeze it and jointly optimize the inference network and stochastic head—allowing uncertainty-aware fine-tuning without catastrophic forgetting of segmentation performance.
What the improved AI system can do:
-
Silent failure detection: Automatically flag inputs from underrepresented subgroups (e.g., infants in pediatric cohorts) by producing elevated uncertainty, even when segmentation accuracy is still moderately high (DSC 0.85), enabling clinicians to know when to trust or reject predictions.
-
Zero-shot OOD robustness: Deploy on new populations without retraining or additional labeled data, with uncertainty that reliably tracks error (NCC > 0.6 after calibration), outperforming deep ensembles and PHiSeg by 10–30% in uncertainty quality.
-
Reliable downstream clinical decisions: For tasks like ejection fraction estimation, achieve lower mean absolute error with substantially reduced variance on OOD subgroups, and improve detection of cardiac malfunction (AUROC 0.94 vs 0.90 for ensembles) by leveraging uncertainty to down-weight unreliable predictions.
-
Post-hoc calibration compatibility: Maintain superior uncertainty ranking even after temperature scaling, meaning the advantage is structural (not just better-scaled confidence) and cannot be replicated by simple recalibration of existing methods.
-
Computationally efficient uncertainty: Provide per-pixel uncertainty from a single forward pass with a tiny stochastic head, avoiding the 10× inference cost of deep ensembles, while still matching or exceeding their uncertainty quality on OOD data.
Abstract
Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in training data. A common case is pediatric care, where models trained on adult cohorts can silently under-perform on children with no indication that something has gone wrong. As retraining with labeled pediatric data is often infeasible, detecting such failures at inference time is a critical clinical need. Building on the VIDS (Variational Inference under Distribution Shifts) framework, we introduce VIDS-Seg, which applies amortized variational inference over a lightweight prediction head to make this adaptive, OOD-aware prior tractable for dense image segmentation. We evaluate VIDS-Seg on left ventricular segmentation in echocardiography, a setting where pediatric anatomy differs systematically from the adult population most segmentation models are trained on, training on an adult cohort (EchoNet-Dynamic) and evaluating zero-shot on a pediatric cohort (EchoNet-Pediatric). Across all age strata, VIDS-Seg matches competitive baselines in segmentation accuracy while producing substantially higher spatial correspondence between predicted uncertainty and segmentation error, an advantage that persists even after applying temperature scaling to all baselines. Downstream, it yields more accurate and stable ejection fraction estimates and more reliable detection of cardiac malfunction in the infant subgroup. Our results indicate that OOD-aware uncertainty quantification can serve as a practical safety layer for deployed segmentation models, enabling detection of silent failures in underrepresented subgroups without retraining or additional labeled data.
Sources
- Out-of-distribution Detection in Medical Image Analysis: A survey
- Improving Uncertainty-based Out-of-Distribution Detection for Medical Image Segmentation
- Quantifying Uncertainty in the Presence of Distribution Shifts
- A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
- Distribution-Free, Risk-Controlling Prediction Sets
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models