Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification
Max A. Nelson, Eminenur Sen Tasci, Zhixiang Wang, Zongwei Zhou, Halil Ertugrul Aktas, Andrea M. Bejar, Elif Keles, Ziliang Hong, Sıtkı Safa Taflan, Muhammed Enes Tasci, Frank H. Miller, Michael B. Wallace, Rajesh N. Keswani, Gorkem Durak, Ulas Bagci
Northwestern University · Johns Hopkins University · Istanbul University · Mayo Clinic Florida
cs.CV, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: Accepted at the International Workshop on Machine Learning in Medical Imaging (MLMI 2026), held in conjunction with MICCAI 2026. This is the authors' accepted manuscript; the final version will appear in Springer Lecture Notes in Computer Science (LNCS). 11 pages, 3 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces cUPMI (calibrated Upstream Probabilistic Meta-Imputation), a class-conditional Gaussian augmentation method for regularizing level-1 stacking ensembles in multimodal IPMN
Terminology
Summary
The paper introduces cUPMI (calibrated Upstream Probabilistic Meta-Imputation), a class-conditional Gaussian augmentation method for regularizing level-1 stacking ensembles in multimodal IPMN (intraductal papillary mucinous neoplasm) risk stratification. The study uses a multi-center cohort of 678 patients from the Cyst-X collection, with paired T1W/T2W MRI sequences and whole-organ (n=629) and sub-regional (n=616) pancreas segmentations. The pipeline stacks base classifiers—radiomics random forests and 2.5D ImageNet-pretrained ResNet-18 deep learning streams—across 16 streams defined by representation (radiomics or CNN), sequence (T1W/T2W), and region (whole, head, body, tail). The meta-features are per-class log-probabilities concatenated across streams, and cUPMI fits one class-conditional Gaussian per class with a shared pooled covariance (ridge λ=10−4) in log-probability space, then injects synthetic samples (ratio ρ selected by inner cross-validation) into the training fold of the level-1 combiner. Three combiners are evaluated: L2-regularized logistic regression, random forest (200 trees, max depth 4), and XGBoost (200 shallow trees, max depth 3, learning rate 0.05). Binary classification uses AUC; ordinal 3-class (no<low<high) uses quadratic-weighted Cohen's κ (QWK) with macro one-vs-rest AUC.
Key results: cUPMI provides no benefit to the already-regularized L2 logistic stack in binary tasks (∆+0.0005 AUC) but consistently improves higher-capacity tree combiners in binary and radiomics-only settings—RF gains +0.015 and XGBoost +0.024 binary AUC, positive in all seeds. In ordinal classification, cUPMI's clearest benefit is for XGBoost on an 8-stream radiomics task (+0.022 QWK, positive in all seeds), while RF (+0.004) and LR (+0.007) show minimal gains. In the fused DL+radiomics setting, cUPMI fails to improve RF (∆QWK −0.003) and slightly hurts LR (−0.007), with XGBoost retaining a negligible +0.005 gain. The strongest overall model is a fold-locked RF stack over fused DL+radiomics without cUPMI, achieving QWK 0.595 (95% CI [0.54, 0.64]) and binary AUC 0.839, surpassing radiomics (QWK 0.514, AUC 0.813), 2.5D ResNet-18 (QWK 0.497, AUC 0.784), and 3D DenseNet-121 (QWK 0.452, AUC 0.774) baselines. Errors concentrate on adjacent grades with near-zero two-step confusions (0.02/0.04). The paper concludes that cUPMI regularizes higher-capacity tree combiners, and that careful fusion of whole-organ and sub-regional radiomics with deep-learning streams outperforms heavier single architectures on small ordinal cohorts. External validation and full multiclass calibration are noted as future work.
Improvements for AI systems
Improvements to AI systems:
-
Add class-conditional Gaussian augmentation with pooled covariance to tree-based ensemble meta-learners (random forests, XGBoost) in stacking pipelines. This injects synthetic per-class log-probability samples during training, which regularizes high-capacity combiners and yields +0.015 to +0.024 AUC gains in binary tasks, with consistent positive results across all random seeds.
-
Use per-class log-probabilities as meta-features instead of raw predictions or hard labels when stacking heterogeneous base models (radiomics + deep learning). This preserves calibration information and enables the Gaussian augmentation to operate in a well-behaved, bounded space.
-
Apply cUPMI selectively—only for tree-based combiners, not for already-regularized linear models (e.g., L2 logistic regression). The method provides negligible benefit (+0.0005 AUC) for linear stacks, so computational cost can be avoided there.
-
For ordinal classification (e.g., no<low<high), use quadratic-weighted Cohen’s κ as the optimization target and evaluate cUPMI on radiomics-only streams—the clearest gain is +0.022 QWK for XGBoost on an 8-stream radiomics task, positive across all seeds. Avoid applying cUPMI to fused deep-learning+radiomics ordinal stacks, where it can slightly hurt (RF −0.003, LR −0.007).
-
Fuse whole-organ and sub-regional (head/body/tail) radiomics with 2.5D ImageNet-pretrained ResNet-18 streams using a fold-locked random forest combiner (without cUPMI) for the strongest overall performance: QWK 0.595, binary AUC 0.839, surpassing heavier 3D DenseNet-121 (QWK 0.452, AUC 0.774) on small ordinal cohorts.
-
Select the augmentation ratio ρ via inner cross-validation on the training fold, rather than fixing it globally, to adapt the synthetic-sample injection to the specific combiner and task.
What the improved AI system can do:
-
Robustly stratify IPMN risk (no/low/high) with higher accuracy and better ordinal consistency (near-zero two-step misclassifications, 0.02/0.04) than single heavy architectures.
-
Automatically choose whether to apply cUPMI based on combiner type (tree vs. linear) and modality fusion (radiomics-only vs. fused), saving compute while maximizing AUC/QWK.
-
Generalize better on small medical imaging cohorts (n≈600–700) by regularizing high-capacity meta-learners, reducing overfitting to training folds.
-
Provide calibrated per-class probabilities from the stacked ensemble, enabling downstream decision support with confidence thresholds.
-
Extend to other multimodal ordinal classification tasks (e.g., disease staging, severity grading) where paired imaging sequences and sub-regional segmentations are available.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models