DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis".
Jane: Medical image segmentation models often perform unevenly across different patient subgroups, and existing fairness methods frequently fail by treating each subgroup as internally homogeneous,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about who wrote this, Jane. The paper "DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis" features Yiqi Tian and Sangjoon Park as the main researchers, with contributions spread across a few different institutions.
Jane: It shows a collaborative effort between groups like the Center for Advanced Medical Computing and Analysis at Massachusetts General Hospital and the Department of Industrial Engineering at the University of Pittsburgh, which suggests this work draws from diverse expertise.
Lu: The team structure reflects how complex these fairness problems are; they need input from both clinical environments like radiation oncology and computational engineering to build a comprehensive framework.
Meng: It's interesting that you see those affiliations; it means their approach isn't just theoretical, but it has roots in actual medical computing and real-world application needs.
Lalam: I think the diversity of the authors hints at the complexity of medical data itself, which is inherently heterogeneous across many different patient profiles.
The paper's summary: Tom: Now, to get into what they actually propose, the DuetMoE paper summarizes their core idea as a joint framework that addresses both inter-subgroup differences and intra-subgroup variation in medical image segmentation models.
Jane: Essentially, they are trying to solve the problem where current fairness methods only look at improving average subgroup performance, which often ignores those hard cases buried within each group.
Lu: They propose FairDRO as the concrete method that marries subgroup-aware distributionally robust optimization with a mixture-of-experts architecture to capture this two-pronged approach of adaptation and robustness.
Meng: So, the summary suggests they are using this combination to ensure the model handles both variations in what's present across groups and variations in how hard the specific images are within a group.
Lalam: It seems they’ve framed it as a way to stop those difficult, high-loss samples from getting washed out by the subgroup average, which is a really smart way to think about stabilizing performance.
The paper's improvements: Tom: Regarding the improvements they propose in "DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis," the main advancement is their concrete implementation of FairDRO which couples subgroup-aware mixture-of-experts with a subgroup-conditioned distributionally robust objective.
Jane: This means they are using the dMoE component to adjust how features are learned across different subgroups, while the DRO part focuses on emphasizing high-loss samples within each specific group.
Lu: The paper details how the dMoE component captures inter-subgroup heterogeneity by adapting intermediate features using a subgroup-conditioned mixture of expert transformations, which is a key difference from just relying on loss weights.
Meng: From an engineering view, this means the routing mechanism for those experts gets informed by the patient's subgroup attribute, making the feature extraction itself contextually aware of that group.
Lalam: The intra-group robustness part is handled by defining a subgroup-specific ambiguity set Ug(Pbg), which lets them evaluate risk based on distributions close to that specific subgroup’s data, prioritizing those hard cases.
Conclusion: Tom: So, wrapping up the DuetMoE paper, the main conclusion is that their FairDRO mechanism successfully couples inter-subgroup adaptation with intra-subgroup robustness to create a model where per-patient prediction becomes tighter across various subgroups.
Jane: They showed that this approach improves worst-case subgroup performance significantly on datasets like Harvard-FairSeg and HAM10000, demonstrating that the system makes fewer patients receive poor segmentation results.
Lu: The implications are huge because it shows a practical way to use representation learning and robust optimization together to solve the specific problem of intra-group hidden failure in medical image analysis.
Meng: For me, the practical impact is seeing how this could translate into more reliable clinical tools where we can trust that the contouring or delineation for a difficult case won't be missed just because it belongs to a smaller subgroup.
Lalam: I think this work has major implications for the future of medical AI because it sets a direction for training systems that are designed not just to perform well generally, but to be dependable on an individual patient level across all demographic and clinical settings.
Yiqi Tian, Sangjoon Park, Bo Zeng, Pengfei Jin, Yujin Oh, Quanzheng Li
Center for Advanced Medical Computing and Analysis, Massachusetts General Hospital and Harvard Medical School · Department of Industrial Engineering, University of Pittsburgh Department of Radiation Oncology, College of Medicine, Yonsei University Department of Biomedical Systems Informatics, College of Medicine, Yonsei University
cs.CV, cs.AI
Submitted: 2026-05-11
Updated: 2026-09-29
Comments: 24 pages, 7 figures, 12 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: Medical image segmentation models often perform unevenly across different patient subgroups, and existing fairness methods frequently fail by treating each subgroup as internally homogeneous, which
Key concepts
- Inter-Subgroup Axis (Representation Learning)
- This axis focuses on adapting the model's internal features to capture differences between various patient subgroups. It uses a subgroup-conditioned mixture of expert transformations, allowing different experts to learn distinct patterns related to demographic and clinical variations, rather than relying solely on loss weights.
- Intra-Subgroup Axis (Robust Loss Aggregation)
- This axis ensures model robustness within each specific subgroup by considering the worst-case distribution of samples. It defines a subgroup-specific ambiguity set that captures distributions close to the actual subgroup data, making the model sensitive to difficult, hard cases within that group.
- FairDRO Objective
- The final objective function combines both axes by aggregating robust risks across all subgroups. By setting weights uniformly, the framework ensures that subgroup information guides expert routing while robustness is enforced by considering the worst-case loss distribution within each specific subgroup.
Terminology
Summary
Medical image segmentation models often perform unevenly across different patient subgroups, and existing fairness methods frequently fail by treating each subgroup as internally homogeneous, which obscures difficult cases within those groups. This study proposes DuetFair, a dual-axis fairness framework implemented via FairDRO, to jointly address inter-subgroup adaptation and intra-subgroup robustness in medical image segmentation. The proposed mechanism aims to ensure that model outputs are reliable for individual patients across various demographic and clinical subgroups by accounting for both distribution shifts between groups and hard samples within those groups.
The gist
FairDRO combines subgroup-aware distributionally robust optimization (DRO) loss aggregation with a subgroup-conditioned mixture-of-experts (dMoE) representation to simultaneously capture inter-subgroup heterogeneity and intra-subgroup variation in medical image segmentation.
How it works: The DuetFair Mechanism
The DuetFair mechanism views subgroup fairness as a joint problem of inter-group adaptation and intra-group robustness.
This is realized through two complementary axes:
-
Inter-Subgroup Axis (Representation Learning): FairDRO uses dMoE to address inter-subgroup variation at the representation level. Specifically, it adapts intermediate features using a
subgroup-conditioned mixture of expert transformations
defined by Equation (2), which allows different experts to captureheterogeneous demographic and clinical patterns.
This placesinter-subgroup adaptation in the representation module, rather than relying only on loss weights to capture the group-wise differences.
-
Intra-Subgroup Axis (Robust Loss Aggregation): The intra-subgroup axis is handled through a distributionally robust view of subgroup risk. Instead of evaluating loss under a uniform empirical mean, FairDRO defines a
subgroup-specific ambiguity set Ug(Pbg) containing distributions close to Pbg.
The robust risk within subgroup g is defined as the supremum over this set:
Rrob g(θ, ϕ):= sup Qg∈Ug(Pbg) E(x,y)∼Qg lf dMoE θ,ϕ (x, g), y. (3)
This formulation makes the robust risk sensitive to high-loss samples within subgroup g,
ensuring that hard individual cases are prioritized relative to their own subgroup while the subgroup structure remains explicit.
How it works: The FairDRO Objective
The final objective couples these two axes through a weighted aggregation of the robust risks. The overall objective is formulated as:
min θ,ϕ LFairDRO(θ, ϕ) = X g∈G wgRrob g(θ, ϕ), (4)
Here, Rrob g(θ, ϕ) is the within-subgroup robust risk evaluated using the dMoE predictor. The weight wg aggregates subgroup-level robust risks. In the implementation described in Equation (4), we set wg = 1/G, so FairDRO uniformly aggregates subgroup risks without manually upweighting any subgroup.
This design ensures that subgroup information guides dMoE routing, while robustness is imposed through the worst-case distribution over samples within each subgroup.
Evaluation and Results
FairDRO was evaluated on three medical image segmentation benchmarks: Harvard-FairSeg, HAM10000, and an in-house 3D radiotherapy target cohort. The results demonstrate that FairDRO achieves the best equity-scaled performance on Harvard-FairSeg
and improves worst-case subgroup performance on HAM10000 under both age- and race-based grouping schemes.
On the 3D radiotherapy target cohort, FairDRO further improves worst-group Dice by 3.5 points (↑ 6.0%) under the tumor-stage grouping and by 4.1 points (↑ 7.4%) under the institution grouping over the strongest baseline.
Qualitative analysis shows that FairDRO produces tighter per-patient prediction
across both minor and major groups, rather than a uniform shift toward a subgroup-mean template.
Limitations and Future Directions
The paper notes two primary limitations:
-
FairDRO may not always produce the largest gain when predefined subgroup labels already explain most of the performance gap, suggesting its value is maximized when
subgroup labels only partially describe patient difficulty.
-
FairDRO relies on subgroup attributes (e.g., race, age, tumor T-stage) during training and inference; future work plans to extend DuetFair on the
attribute-aware side
through continuous conditioning and on theattribute-agnostic side
by developing methods such aslatent-attribute inference, attribute imputation, and self-supervised subgroup discovery.
Additionally, an alternative objective variant, FairDRO-Penalty (Equation 16), was explored but showed that direct coupling of robustness with the objective leads to more stable improvements than simply appending a DRO term to the empirical loss.
Improvements for AI systems
As a fastidious researcher, I have analyzed the proposed DuetFair mechanism and FairDRO framework. The core innovation lies in jointly addressing inter-subgroup heterogeneity (via subgroup-conditioned distributionally robust optimization, dMoE) and intra-subgroup hidden failures (via the distributionally robust loss term).
Here are the specific improvements and capabilities that can be derived from this research, targeted at enhancing medical image segmentation systems:
The proposed system, FairDRO, is a dual-axis fairness framework that integrates subgroup-aware representation learning with within-subgroup robustness. It fundamentally shifts the training objective from simple average performance to a mechanism that prioritizes high-loss samples within each subgroup while adapting representations across different subgroups.
Here are the specific improvements and capabilities this system enables:
-
The ability to mitigate
intra-group hidden failures.
-
The capability to achieve superior worst-case subgroup performance under heterogeneous distributions.
-
Enhanced reliability for minority or difficult patient subgroups without sacrificing overall population average accuracy significantly (as demonstrated in Table 6).
Specific, detailed improvements and resulting capabilities:
-
An AI system can be trained to perform medical image segmentation (e.g., tumor delineation, organ boundary tracing) where the performance is not uniformly good across different demographic or clinical subgroups (e.g., race, age groups, or institutional settings).
-
The system will explicitly learn subgroup-aware feature representations using a Distributionally Robust Mixture-of-Experts (dMoE). This means the model can dynamically route input features to the most appropriate
expert
based on the patient's subgroup attribute (e.g., routing an image from a specific age group to an expert trained predominantly on similar anatomical variations). -
The system will incorporate a distributionally robust loss aggregation that acts as an intra-subgroup robustness mechanism. This mechanism prevents high-loss, difficult samples within a specific subgroup (which would otherwise be
washed out
by the subgroup average) from being ignored during training. The model is forced to learn representations that are stable and accurate for these hard cases, even when they represent a small subset of the total data. -
The resulting segmentation output will exhibit tighter per-patient prediction consistency across subgroups rather than just shifting toward a generalized subgroup mean. This means the system produces highly reliable results for individual patients, reducing the risk associated with misdiagnosis in vulnerable patient populations (e.g., ensuring accurate radiotherapy target contouring for rare T1 or T4 tumors, or precise optic cup delineation for specific ethnic groups).
-
In clinical deployment, this system can be configured to maintain high average performance across the entire population while simultaneously guaranteeing that the performance degradation experienced by any single subgroup is minimized (as measured by Equity-Scaled metrics like ES-Dice and ES-IoU).
This improved AI system moves beyond simple group balancing or loss penalization. It uses a sophisticated, two-pronged approach—adapting what the model sees
(representation learning) and adapting how the model learns from hard examples
(robust loss aggregation)—to ensure that no difficult patient case is hidden by the statistical average of their subgroup.
Abstract
As medical AI expands across diverse healthcare settings worldwide, equitable performance across patient populations is becoming essential to trustworthy clinical use. Fairness in medical image analysis is often evaluated through average performance across predefined subgroups, yet similar subgroup averages can conceal substantial variation among individual patients. Therefore, a reliable medical AI requires addressing two complementary objectives: inter-subgroup fairness, which reduces performance disparities across groups, and intra-subgroup robustness, which protects poorly served patients within each group. To jointly address these objectives, we propose DuetMoE, a subgroup-aware mixture-of-experts framework that couples group-level adaptation with patient-specific clinical guidance, enabling more reliable medical image analysis for individual patients. For settings without linked clinical records, we further introduce DuetMoE+, which retains subgroup-routed experts and incorporates a KL-constrained distributionally robust objective to address intra-subgroup robustness. We evaluate our methods on PI-CAI, radiotherapy, and Harvard-FairSeg. Across four evaluation settings, our methods lead in overall or equity-scaled performance, reduce the inter-subgroup mean Dice gap by up to 30.7%, and raise intra-subgroup 25th-percentile Dice by up to 11.5 points over the strongest reported baselines.
Sources
- Customized Segment Anything Model for Medical Image Segmentation
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Tilted Empirical Risk Minimization
- Mixture of Multicenter Experts in Multimodal AI for Debiased Radiotherapy Target Delineation
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Decoupled Weight Decay Regularization
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models