CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion

arXiv:2603.16551 · cs.CV, cs.AI · Submitted 2026-03-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion".

Tom: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we’ve established that the paper is titled "CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion," which basically tells us they are trying to generate high-quality medical images for any combination of demographics, even ones they haven't seen before.

Jane: Exactly, Tom. They are using a hierarchical approach within diffusion models to make sure the quality isn't just good on average, but high for every single demographic intersection they aim to represent.

Lu: What I find compelling is their framework because it treats demographic identity not as a single feature you can just optimize, but as something that can be broken down into smaller, manageable pieces that interact in specific ways.

Meng: That sounds complex when you think about implementation; if we have to build this hierarchical structure, we're adding a lot of components on top of our existing diffusion pipeline. What does the actual architecture look like?

Lalam: From an engineering standpoint, if this works as described, it means our future generative AI won't just be good at what it was trained on; it will have the ability to synthesize novel and fair examples entirely.

The paper's summary: Tom: So, to summarize what they’re proposing with CompDiff, they are addressing the "imbalanced generator problem," which is when diffusion models produce lower quality images for rare subgroups because those groups weren't well-represented in the training set.

Jane: That’s a clear way to put it; while other methods try to fix this by tweaking how we optimize things during training, CompDiff focuses on fixing the representation itself so that rare intersections can be composed from simpler, well-learned parts.

Lu: The core idea is that demographic identity is compositional, meaning a rare intersection isn't random noise; it can be built by combining known single attributes and known pairwise interactions.

Meng: So instead of relying on massive amounts of data for every single subgroup, they are using these learned relationships to predict what an unseen intersection should look like based on the parts we already understand. That’s a smart way to handle data scarcity.

Lalam: It suggests that our generative AI can learn general rules about human anatomy and demographics that apply across different combinations, which is a huge step toward making these systems truly versatile and fair in practice.

The paper's improvements: Tom: Now let’s talk about the specific technical improvements they’ve made, because the methodology is where this paper really shines—they introduce a Hierarchical Conditioner Network, or HCN, to decompose demographic conditioning into single-attribute, pairwise, and composed representations.

Jane: That decomposition is key; they embed attributes like age or sex separately as "grandparents," model their non-additive relationships as "parents," and then combine those interactions into a final "child" representation.

Lu: The way they model the parents using dedicated MLPs to capture non-additive relationships, like how age might interact with bone density, shows a deep understanding of how these factors influence anatomy.

Meng: I see them defining the final demographic representation as something that gets mapped to a diagonal Gaussian and projected into the cross-attention dimension; that’s the practical mechanism for feeding this structured information into the main generation network.

Lalam: This structural regularization, including terms like Compositional Consistency Loss, seems designed to keep the model stable while still allowing those complex interactions to form without completely breaking training.

Conclusion: Tom: So, to wrap things up on "CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion," they have shown that this hierarchical approach works well across chest X-rays and fundus images.

Jane: They achieved the best scores in terms of image quality metrics, like FID, while simultaneously showing the lowest Equity-Scaled FID across sex, race, and age groups on both modalities.

Lu: The ablation studies really support their architecture; they showed that only the hierarchical composition structure could recover high demographic accuracy for attributes like sex at zero point nine nine and race at zero point nine six while keeping image quality reasonably high, as seen in the FID of seventy-five point five for that configuration.

Meng: From an engineering standpoint, it’s encouraging to see that this structural approach is more effective than trying to just compete for a shared text prompt budget when dealing with these complex demographic controls.

Lalam: Ultimately, this work suggests we can move toward generative AI systems that are not just accurate on average, but demonstrably equitable across all patient groups by composing knowledge instead of just memorizing training examples.

Institute of Data Science, Faculty of Science and Engineering, Maastricht University · Department of Advanced Computing Sciences, Faculty of Science and Engineering, Maastricht University · VITO

cs.CV, cs.AI

Submitted: 2026-03-17

Updated: 2026-09-28

Comments: v4: substantially revised version (new title, reader study, additional co-authors). 38 pages main text + 25 pages supplement

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across

Key concepts

Imbalanced Generator Problem
Diffusion models trained on uneven data struggle to generate high-quality images for rare subgroups or combinations of attributes not seen during training. CompDiff solves this by modeling demographics compositionally, enabling generalization even for these previously unseen intersections.
Hierarchical Conditioner Network (HCN)
This network breaks down demographic information into three levels: single attributes (like age), pairwise interactions (how two attributes relate), and a final composed representation. This structure allows the model to capture complex, non-additive relationships between different demographic factors.
Compositional Generalization
The core idea is that a rare intersection of traits can be accurately predicted by combining embeddings learned from individual traits and their pairwise interactions. Instead of needing specific training data for every combination, the model learns the rules to 'compose' new demographic identities.
Equity-Scaled FID (ES-FID)
This metric measures image quality specifically across different demographic subgroups. CompDiff achieved the lowest ES-FID, demonstrating superior performance in generating high-quality images that are equally fair and representative of all tested groups.

Terminology

Summary

Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups. The gist is that CompDiff, a hierarchical compositional diffusion framework, addresses the imbalanced generator problem by decomposing demographic conditioning into single-attribute, pairwise, and composed representations to enable compositional generalization to rare or unseen intersections.

The core problem addressed is the imbalanced generator problem.

Diffusion models trained on imbalanced data inherit these imbalances, degrading synthesis for rare subgroups and struggling with intersections absent from training—the imbalanced generator problem. While remedies like loss reweighting operate at the optimization level and provide limited benefit when training signal is scarce or absent, CompDiff addresses this at the representation level. The paper posits that demographic identity is compositional: a rare intersection can be composed from well-learned single-attribute embeddings and moderately learned pairwise interactions, enabling generalization even to combinations entirely absent from training.

The proposed solution involves the Hierarchical Conditioner Network (HCN).

CompDiff introduces a dedicated Hierarchical Conditioner Network (HCN) that decomposes demographic conditioning into single-attribute, pairwise, and composed representations. This structure produces a demographic token concatenated with CLIP embeddings as cross-attention context. The hierarchy is structured as follows:

  1. Single-Attribute Embeddings (“grandparents”): Each attribute (age, sex, race) is embedded into a shared latent space to produce embeddings like ea, es, er.

  2. Pairwise Interactions (“parents”): All pairwise interactions are modeled using dedicated MLPs to capture non-additive relationships between attributes (e.g., age and bone density).

  3. Full Composition (“child”): The final demographic representation is obtained by combining pairwise interactions through an MLP g(·), resulting in hdemo, which is then mapped to a diagonal Gaussian to sample the latent vector z, which is projected to the cross-attention dimension c.

The training objective incorporates structural regularization and auxiliary supervision.

The model is trained end-to-end with a total loss defined as:

L = Ldiff + λcompLcomp + λauxLaux + λKLLKL.

(1) Diffusion Loss (Ldiff):

Ldiff = Ex0,ϵ,t∥ϵ − ϵθ(xt, t, Ecombined)∥2

The regularization terms are:

(2) KL Term (LKL):

LKL = E[KL(N(µ, σ2 I) N(0, I))] to regularize the variational demographic latent toward a standard normal.

(3) Compositional Consistency Term (Lcomp):

Lcomp = 1 − cos(hdemo, eage + esex + erace), which acts as a soft anchor that stabilizes training toward a simple additive baseline while still allowing non-additive interactions.

(4) Auxiliary Loss (Laux):

This term ensures demographic information survives projection into the cross-attention space by applying it directly on the projected token c: Laux = CE(ˆyage, yage) + CE(ˆysex, ysex) + CE(ˆyrace, yrace).

Evaluation demonstrates superior performance across quality and fairness.

CompDiff was evaluated on chest X-rays (MIMIC-CXR) and fundus images (FairGenMed). Key findings include:

(1) Image Quality:

CompDiff achieved the best FID on both modalities (64.3 chest, 54.6 fundus), outperforming FairDiffusion across image quality metrics.

(2) Subgroup Equity:

CompDiff achieves the lowest Equity-Scaled FID (ES-FID) across sex, race, and age on both modalities. In zero-shot generalization to held-out demographic subgroups, CompDiff outperforms both baseline (B) and FairDiffusion (FD) on all five held-out intersections, lowering FID by up to 21%.

(3) Downstream Utility:

Downstream classifiers trained on CompDiff data show improved AUROC and reduced demographic bias. On chest X-rays, the model achieves a higher mean AUC (0.82 vs 0.74) and lower underdiagnosis rates (0.40 vs 0.46).

Ablation studies confirm the necessity of the hierarchical structure.

Ablations validate that architectural inductive bias is critical for success:

(1) Architectural Necessity:

Among tested architectures, only hierarchical composition (HCN) succeeded in recovering high demographic accuracy (sex 0.99, race 0.96) while maintaining good image quality (FID 75.5). Flat MLP Encoders failed to recover control (sex 0.

Improvements for AI systems

Based on the CompDiff framework described in this paper, here are specific improvements that can be made to existing medical image generation and diagnostic AI systems:


)Improved AI System Capabilities:

  1. [Zero-Shot Intersectional Image Generation]: The system can generate high-fidelity medical images (e.g., chest X-rays, fundus images) for demographic intersections that were entirely absent from the original training data (e.g., 80+ Asian female).

  2. [Demographically Fair Synthesis]: The generation process is explicitly designed to produce equally high-quality images across all demographic subgroups, minimizing quality disparities between groups (as measured by ES-FID).

  3. [Compositional Generalization in Clinical Context]: By decomposing demographic conditioning into a hierarchical structure (single attributes, pairwise interactions), the system can synthesize realistic anatomical features even for unseen intersections by composing known components.

  4. [Improved Downstream Diagnostic Fairness]: When used to train diagnostic classifiers, the resulting model exhibits reduced demographic bias (lower Difference in Equalized Odds) and improved overall diagnostic performance (higher ES-AUC), leading to fairer predictions across different patient populations.

  5. [Controllable Representation Learning]: The system allows researchers to maintain high demographic accuracy (e.g., 99% accuracy on sex/race) while simultaneously improving generative quality (FID reduction) by employing the hierarchical conditioning structure instead of competing in a shared text prompt space.

)Specific Implementation Improvements:

  1. [Architectural Integration for Generation]: Replace standard text encoders or conditioning methods with the proposed Hierarchical Conditioner Network (HCN). This involves replacing the single demographic token concatenation with the combined context vector:

Instead of: [Clinical Text Embeddings] + [Demographic Token competing in CLIP budget]

Use: [Clinical Text Embeddings (CLIP)] + [CompDiff Demographic Token 'c' derived from HCN, concatenated as Ecombined = [Etext, c]] consumed by the UNet via cross-attention.

  1. [Hierarchical Conditioning Module]: Implement the specific decomposition within the HCN for demographic attributes (age, sex, race):

a) Embed individual attributes into shared latent spaces: Create embeddings for age, sex, and race that are mapped to a common dimension (e.g., 256).

b) Model non-additive interactions: Introduce dedicated Multi-Layer Perceptrons (MLPs) to explicitly model pairwise relationships between these embeddings (e.g., calculating interaction features like age-bone density or sex-cardiotoracic ratio).

c) Compose the final representation: Use a final MLP to combine these pairwise interaction features into the full demographic token, which is then projected to the cross-attention dimension 'c'.

  1. [Structured Training Objective]: Modify the training loss function to enforce both structural stability and demographic informativeness:

a) Diffusion Loss: Maintain the standard diffusion loss (Ldiff).

b) Compositional Consistency Loss (Lcomp): Add a soft anchor term, such as Lcomp = 1−cos(hdemo, eage+esex+erace), to stabilize training towards simple additive baselines while allowing non-additive interactions.

c) Auxiliary Classification Loss (Laux): Apply auxiliary classification directly on the projected token 'c' (not the latent mean µ) using Cross-Entropy loss for each attribute (age, sex, race). This ensures the final representation utilized by the UNet is explicitly demographically informative.

  1. [Hyperparameter Tuning for Stability]: Use a carefully balanced regularization strategy during training:

a) Variational Regularization: Apply KL divergence loss (LKL) to keep the demographic latent distribution near a standard normal, ensuring smooth sampling.

b) Loss Weighting: Carefully tune the weighting parameters (λcomp and λaux). Empirical evidence suggests balancing these terms is critical to achieving the best trade-off between image quality (FID) and subgroup fairness (ES-FID).

Abstract

Medical image generators trained on imbalanced data can fail at demographic intersections absent from training. We introduce CompDiff, which encodes age, sex and race separately and composes supervised demographic tokens alongside clinical text. Across chest radiographs and fundus images, CompDiff improves overall and subgroup fidelity relative to prompt conditioning (RoentGen-v2) and loss reweighting (FairDiffusion). It generalises in zero-shot generation to 16 chest X-ray intersections excluded from training, achieving the lowest mean FID-RadImageNet in every intersection. In a blinded reader study of these unseen intersections, two radiologists gave CompDiff the highest mean scores among generators for anatomical realism and agreement with the clinical impression, and selected its images most often as the most realistic. Pretraining with CompDiff images improved downstream classification, while CompDiff audit cohorts reduced estimation error on rare intersections. These findings support compositional demographic conditioning for extending medical image synthesis to underserved populations. Code: https://github.com/mahmoudibrahim98/CompDiff

Sources

Related papers