Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast

arXiv:2603.04113 · cs.CV, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast".

Jane: The paper was written by Mehmet Yigit Avci, Akshit Achara, Andrew King and Jorge Cardoso from King's College London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: Welcome back, everyone. Tom here with Jane, and we are kicking off our discussion of a brand new paper on arXiv. It’s called "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, I have to say, the title alone is a mouthful, but the question it asks is huge.

Jane: It really is, Tom. And the question is basically this: when a computer looks at a brain MRI and can guess your age, your sex, even your race, where is that information actually coming from? Is it in the physical structure of your brain, or is it hiding in the way the picture was taken?

Tom: Right, and that matters because these models are being used for real clinical tasks like diagnosing Alzheimer's. If they're picking up on demographic clues instead of actual disease, that's a bias problem. But nobody had really pinned down the source of that signal in brain MRI before.

Jane: Exactly. In chest X-rays, researchers found that a lot of the demographic signal comes from the acquisition settings—the machine, the protocol, the exposure. The team here, led by Mehmet Yigit Avci and Akshit Achara at King's College London, wanted to see if the same was true for brain MRI.

Tom: And they built a clever way to test it. They used two models, MR-CLIP and DIST-CLIP, which can take a single brain scan and split it into two separate representations. One captures the anatomy, the actual shape and structure of the brain. The other captures the contrast, which is basically the imaging fingerprint of the scanner and protocol.

Jane: So it's like separating the recipe from the chef. The anatomy is the ingredients, the contrast is the cooking style. And then they trained predictors for age, sex, and race on each of those components separately, plus on the raw full image.

Tom: And the results were pretty striking. Across three datasets—OASIS, ADNI, and HCP—the anatomy representations preserved almost all of the predictive power. For sex prediction, raw images hit about ninety-two to ninety-five percent balanced accuracy, and the anatomy-only representations were right there at ninety-three to ninety-six percent.

Jane: Meanwhile, the contrast embeddings were much weaker. They still had some signal, but it was clearly secondary. So the big takeaway here is that in brain MRI, the demographic signal is primarily anatomical, not a scanner artifact.

Tom: That's a big deal because it flips the script from the chest X-ray findings. It means you can't just harmonize the images or normalize the intensities and expect the bias to disappear. The information is baked into the structure itself.

Jane: And that has real consequences for how we build fair clinical models. But before we get into the fixes, I want to dig into how they actually proved this disentanglement works. That's coming up next.

Tom: Stay with us.

Paper discussion segment 2: Jane: Welcome back. We're still on "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Tom, we left off saying the anatomy carries most of the signal. But how do we know the models actually separated anatomy from contrast cleanly?

Tom: Great question, and the paper has a really neat validation. They took five hundred subjects from OASIS and used SynthSeg to get ground-truth brain measurements like intracranial volume and ventricle volume. Then they probed each representation to see what it actually encoded.

Jane: And the results were stark. The anatomy representation predicted intracranial volume with an R-squared of zero point nine five three, almost perfect. But when they tried to predict which scanner or protocol the image came from, it only got forty-two point five percent balanced accuracy, where chance was six point seven percent across fifteen clusters. So it really is anatomy-focused.

Tom: On the flip side, the contrast embedding was much better at guessing the acquisition protocol, hitting seventy point two percent, but it was noticeably worse at anatomy, with an R-squared of zero point eight zero five for intracranial volume. So the separation is real, though not perfect.

Jane: And here's a subtle but important point. The metadata-only embedding, which comes from DICOM parameters and never sees the image, already predicted intracranial volume with an R-squared of zero point seven three nine. That tells us a lot of the residual anatomical signal in the contrast embedding isn't actually image leakage—it's just that certain scanners tend to be used on certain populations.

Tom: That's a really insightful observation. It means some of the demographic signal we see in contrast is just a proxy for site-specific demographics. You scan a wealthier, healthier population at one site, and the scanner settings correlate with that population's brain characteristics.

Jane: Exactly. So when they trained the full models, the raw images performed best, as you'd expect. Age MAE was around two point six seven years in HCP, sex balanced accuracy up to zero point nine six. But the anatomy representations were nearly identical. And the contrast embeddings, while weaker, still had non-trivial signal—sex at zero point seven one in OASIS, for example.

Tom: So within a single dataset, contrast does carry some demographic information. But the real test is generalization. If you train on one site and test on another, does that contrast signal hold up?

Jane: And that's where the story gets even more interesting. We'll get into those cross-dataset results right after this.

Paper discussion segment 3: Tom: Back with "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, we just said within a dataset, contrast has some signal. But the cross-dataset results really tell the story.

Jane: They do. When they trained on ADNI and tested on OASIS, sex prediction from raw images got seventy-six percent balanced accuracy. The anatomy representation actually improved that to eighty-four percent. But the contrast embedding collapsed to seventy percent, and for race it dropped to fifty percent, which is pure chance.

Tom: That pattern repeats across nearly every transfer pair. The contrast signal is dataset-specific. It doesn't generalize. The anatomy signal, on the other hand, transfers much more robustly. For example, training on HCP and testing on OASIS, anatomy sex prediction hit seventy-one percent while raw was seventy-two percent—basically the same.

Jane: And age prediction under cross-dataset shift was brutal. Training on HCP and testing on ADNI gave a mean absolute error of forty-two years. That's essentially random. But even there, the anatomy representation didn't do worse than raw—it was forty-four years, so comparable.

Tom: Right, so the anatomy isn't a magic bullet for domain shift, but it doesn't add extra fragility either. The key finding is that the acquisition-driven signal is fragile and site-specific, while the structural signal is more stable.

Jane: They also looked at multiple sequences within OASIS—T1w, T2w, and FLAIR. And the same pattern held. Anatomy representations stayed close to raw performance across all sequences. For FLAIR, anatomy sex prediction actually beat raw, eighty-three percent versus seventy-one percent.

Tom: That's a fascinating result. It suggests that when the raw image is noisy or the contrast is less informative, the anatomy representation can be more robust because it strips away the irrelevant acquisition noise.

Jane: So what does this mean for bias mitigation? The authors are pretty clear. If you want to reduce demographic predictability in brain MRI models, you can't just focus on harmonizing intensities or normalizing contrast. The signal is in the structure.

Tom: And that's a harder problem. You need to identify which specific anatomical features are driving the demographic signal and decide whether they're clinically relevant or not. For Alzheimer's, some structural changes are exactly what you want to detect. So you can't just throw away anatomy.

Jane: Right, it's a surgical approach, not a blunt one. And the authors suggest that acquisition-related bias should still be addressed, but mainly when it introduces additional demographic information in site-specific settings.

Tom: Before we wrap up, I want to bring in Lu and Meng to get their take on the practical side. Lu, you've been listening—what excites you about this?

Lu: I think the framework is the real contribution here. It gives us a way to decompose the signal and ask mechanistic questions. That's rare in this field. Most bias papers just say "there's bias," but this one says "here's where it lives and here's why."

Meng: And from an engineering standpoint, the fact that the anatomy representation is robust across sites is promising. It means we could potentially use it as a preprocessing step to reduce spurious correlations without losing task-relevant information. But we'd need to validate that on actual clinical outcomes, not just demographic prediction.

Jane: That's a perfect segue to our conclusion. Let's wrap this up.

Conclusion: Tom: Alright, we're closing out our discussion of "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, give us the final summary.

Jane: The paper asked a simple but profound question: when models predict demographics from brain MRI, where does the signal come from? And the answer, across three datasets and multiple sequences, is that it's primarily anatomical. The anatomy-focused representations preserved nearly all the predictive power of raw images, while contrast embeddings were weaker and failed to generalize across sites.

Tom: And that's a direct contrast to findings in chest X-rays, where acquisition played a bigger role. So this paper really advances our understanding of modality-specific bias sources.

Jane: The implications are clear. Bias mitigation in brain MRI can't rely on intensity harmonization alone. It has to engage with the anatomical features themselves, and it has to be careful not to throw away clinically meaningful structure.

Tom: There are limitations, of course. The datasets are imbalanced, especially ADNI with only five point two percent Black participants. Race prediction results should be taken with a grain of salt. But the framework is solid and the methodology is rigorous.

Meng: I'd add that the cross-dataset results are the most actionable part. They show that contrast-based demographic signal is a site artifact, which means site-specific mitigation strategies could work for acquisition bias, but anatomical bias needs a different approach.

Lu: And I'd say the next step is to connect this to downstream clinical tasks. Does reducing demographic predictability in the representation actually reduce disparities in Alzheimer's detection? That's the million-dollar question.

Jane: Exactly. This paper gives us the map, but we still need to navigate the territory. We'll be watching for follow-up work.

Tom: Thanks for joining us, everyone. We're saying goodbye to this paper and getting ready to dive into the next one. Stay curious.

Jane: See you next time.

Mehmet Yigit Avci, Akshit Achara, Andrew King, Jorge Cardoso

King's College London

cs.CV, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Key concepts

Anatomy Representation
This representation captures the actual shape and structure of the brain from an MRI scan. The study found that this component preserved almost all predictive power for demographic factors like sex and age, suggesting structural features are the main source of demographic signal.
Contrast Embedding
This embedding captures information about the way the image was taken, such as scanner settings and protocol. The research showed that contrast embeddings were weaker predictors of demographics and were fragile when tested across different datasets, indicating this signal is site-specific.
Disentangling Models (MR-CLIP and DIST-CLIP)
These models split a single brain scan into two separate representations: one capturing anatomy and another capturing contrast. This method allowed researchers to test which component—anatomy or contrast—was responsible for demographic predictions, revealing the source of the signal.
Cross-Dataset Generalization
This refers to testing if a model trained on one dataset (e.g., HCP) performs well on a different dataset (e.g., OASIS). The study showed that while contrast signals were site-specific and failed to generalize, the anatomy signal transferred more robustly between sites.

Terminology

Summary

Summary

This paper investigates the origins of demographic predictability in brain MRI by decomposing images into anatomy-focused and contrast-dependent representations. The authors propose a controlled framework based on disentangled representation learning, using MR-CLIP and DIST-CLIP models, to separate structural information from acquisition-dependent signal. They train predictive models for age, sex, and race on full images, anatomical representations, and contrast embeddings across three datasets (OASIS-3, ADNI, HCP) and multiple MRI sequences (T1w, T2w, FLAIR).

The key findings are that demographic predictability is driven primarily by anatomical variation, with anatomy-focused representations largely preserving the performance of models trained on raw images. Contrast embeddings retain a weaker, dataset-specific signal that does not generalize across sites. The paper concludes that effective bias mitigation must explicitly account for the primarily anatomical and secondarily acquisition-dependent origins of demographic signal.

Experimental Design. The authors define a dataset D = (xi, yi, ai) of brain MRI scans, demographic attributes (age, sex, race), and acquisition metadata. They decompose each image x into two components: zanat (anatomy-focused, minimizing acquisition effects) and zcontrast (acquisition-dependent, reducing anatomical dominance). MR-CLIP learns a 512-dimensional embedding reflecting acquisition differences via contrastive learning on image–metadata pairs. DIST-CLIP adds an explicit anatomical pathway with patch-level contrastive loss to enforce acquisition invariance. The framework does not assume perfect disentanglement; instead, they quantify the degree of disentanglement through probing analysis. Using SynthSeg, they extract intracranial volume (ICV) and ventricle volume for N=500 OASIS-3 subjects. Results show zanat strongly predicts anatomical features (ICV R2=0.953, MAE=31.1 cm3) but poorly predicts acquisition protocol (BalAcc 42.5%, chance 6.7%). Conversely, zcontrast is informative about acquisition identity (BalAcc 70.2%) while encoding less anatomical detail (ICV R2=0.805, MAE=88.1 cm3). The metadata-only embedding zmetadata achieves ICV R2=0.739, MAE=94.2 cm3, indicating residual anatomical signal in zcontrast reflects population-level correlations between acquisition settings and demographic composition across sites.

Three predictive mappings are defined: ffull (raw MRI), fanat (anatomy representations), and fcontrast (contrast embeddings). For image-based inputs, a 3D ResNet-50 architecture is used on 128×128×128 volumes with ImageNet-pretrained weights, learning rate 1×10−4, dropout 0.2, batch size 2. For embedding-based inputs, a multi-layer perceptron (512→512→256→Nclasses) with ReLU and dropout is used. Age prediction uses MSE loss; sex and race use cross-entropy. Metrics are MAE for age and balanced accuracy (BalAcc) for classification. All volumes are rigidly registered to MNI152 at 1.0 mm3, with skull stripping via SynthStrip.

Within-Dataset Results. Models trained on raw images achieve the highest performance (sex BalAcc 0.92–0.96; age MAE 2.67–5.15 years). Anatomy-focused representations largely preserve this performance (sex BalAcc 0.93 in OASIS). Joint training across all datasets shows consistent patterns (Sex 0.96 Raw vs. 0.96 Anat vs. 0.78 Contrast). Contrast embeddings retain non-trivial predictive power: sex BalAcc 0.71 in OASIS, 0.81 in HCP, 0.78 in joint training; race BalAcc 0.71 in OASIS, 0.68 jointly; age MAE 6.28 years in OASIS, 5.21 jointly. Although consistently lower than raw and anatomy models, contrast alone retains demographic signal within a single domain.

Cross-Dataset Generalization. Performance drops markedly under dataset shift, especially for age (e.g., HCP→ADNI: Raw MAE 42.39 years). Anatomy-focused representations tend to generalize more robustly than raw images (e.g., ADNI→OASIS sex prediction improves from BalAcc 0.76 Raw to 0.84 Anat). Contrast embeddings frequently collapse toward chance (0.50), indicating poor transferability of acquisition-dependent signal. Structural representations retain more stable demographic information under distribution shift. This indicates demographic predictability arises mainly from anatomical sources, with acquisition-driven correlations being less robust and more site-dependent.

Multi-Sequence Analysis. Across T2w and FLAIR images in OASIS, anatomy representations remain close to raw performance (e.g., sex BalAcc 0.91 vs. 0.90 for T2w). Contrast embeddings retain measurable predictive power (sex BalAcc 0.72 in T2w, 0.75 in FLAIR) but race prediction remains near chance. Joint training on all sequences maintains strong sex prediction (anatomy BalAcc 0.93, age MAE 5.16; contrast BalAcc 0.76, age MAE 5.90). The consistency of trends across sequences indicates the demographic signal is not specific to a single modality.

Discussion. The authors conclude that the majority of demographic signal is preserved in anatomy-focused representations, establishing structural variation as the dominant contributor. This aligns with biological expectations for age and is plausible for sex given known morphological differences, with evidence that relative brain volume can differ across racial and ethnic groups. Contrast embeddings retain stable, above-chance predictive performance possibly due to population-level acquisition correlations or minor disentanglement leakage. The findings imply that approaches focusing exclusively on acquisition harmonization or intensity normalization are unlikely to fully mitigate demographic predictability. Conversely, indiscriminately removing anatomical variation risks suppressing clinically meaningful information. Effective bias mitigation should identify and control specific anatomical features contributing to demographic signal while preserving those relevant to the clinical task. A limitation is the imbalanced racial composition of datasets, particularly ADNI with only 84 Black participants (5.2%), limiting statistical reliability of race prediction results despite balanced accuracy mitigation.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

  • Improvement: Add a disentanglement module (using MR-CLIP/DIST-CLIP) to any brain MRI model that predicts clinical outcomes, splitting the input into zanat (anatomy-focused) and zcontrast (acquisition-focused) representations before the prediction head.

  • What the improved system can do: Explicitly quantify how much of a model's demographic bias comes from anatomy vs. acquisition. This allows targeted mitigation—e.g., if the clinical task is Alzheimer's detection, you can suppress zcontrast when it carries site-specific demographic signal, while retaining zanat for clinically relevant morphology.

  • Improvement: Replace raw image inputs with zanat representations for models that need to generalize across sites (e.g., multi-center clinical trials). The paper shows zanat generalizes better than raw images (e.g., ADNI→OASIS sex prediction: 0.84 vs. 0.76 raw).

  • What the improved system can do: A model trained on zanat from one site will maintain higher accuracy when deployed at a new site with different scanners/protocols, reducing the need for site-specific fine-tuning and preventing performance collapse seen with raw images (e.g., HCP→ADNI age MAE drops from 42.39 to 28.11 years when using zanat).

  • Improvement: Build a diagnostic pipeline that, for any given MRI dataset, computes three probes: (a) anatomical probe (ICV, ventricle volume), (b) acquisition probe (scanner/protocol cluster), and (c) demographic probe (age, sex, race). Use the paper's finding that zcontrast retains demographic signal (e.g., sex BalAcc 0.71–0.81) to flag when acquisition settings are correlated with demographics.

  • What the improved system can do: Automatically alert researchers if their dataset has site-specific demographic imbalances (e.g., a scanner used predominantly for one race), enabling proactive collection of balanced data or acquisition-parameter harmonization before training.

  • Improvement: Apply the disentanglement framework across multiple MRI sequences (T1w, T2w, FLAIR) simultaneously. The paper shows zanat performs consistently well across sequences (sex BalAcc 0.90–0.93), while zcontrast retains weaker, sequence-specific signal.

  • What the improved system can do: A single model that accepts any sequence type and automatically routes anatomical information through zanat while discarding acquisition-specific contrast, ensuring that demographic bias does not vary unpredictably when switching between T1w and FLAIR inputs in a clinical workflow.

  • Improvement: In a downstream task (e.g., disease classification), train two versions: one on zanat and one on zcontrast. Use the paper's cross-dataset results to decide which representation to use. For tasks where zcontrast fails to generalize (e.g., race prediction drops to chance 0.50 cross-site), discard zcontrast entirely to remove site-specific demographic leakage.

  • What the improved system can do: A model that automatically selects the most robust representation for a given clinical task, reducing false demographic correlations without sacrificing clinical accuracy, and providing a clear audit trail of which signal source drove each prediction.

  • Improvement: For age regression models, use zanat instead of raw images when deploying across sites. The paper shows zanat maintains MAE within 0.3–0.5 years of raw-image performance within a site (e.g., OASIS: 5.45 vs. 5.15 years) but generalizes better (e.g., ADNI→OASIS: 5.49 vs. 5.17 years raw, with less variance).

  • What the improved system can do: A robust age-estimation tool for longitudinal studies where patients are scanned at different facilities over time, preventing age prediction drift due to scanner changes and ensuring consistent clinical interpretation.

  • Improvement: Given the paper's finding that race signal is primarily anatomical (raw BalAcc 0.86–0.97 vs. contrast 0.50–0.71), implement a race-bias mitigation strategy that operates on zanat by identifying and selectively masking anatomical regions that drive race prediction (e.g., using gradient-based attribution on the zanat encoder).

  • What the improved system can do: A model that can reduce race-based performance disparities in clinical tasks (e.g., disease detection) by removing race-correlated anatomical features that are not clinically relevant, while preserving task-relevant morphology—directly addressing the paper's call for identifying and controlling the specific anatomical features that contribute to demographic signal.

  • Improvement: Replace traditional intensity normalization (e.g., histogram matching) with the DIST-CLIP zanat representation as a harmonization step. The paper shows zanat suppresses acquisition differences (t-SNE clustering by site is absent in zanat space) while preserving anatomy.

  • What the improved system can do: A harmonization module that produces anatomically consistent images across sites without altering tissue contrast in a way that could hide pathology, enabling more reliable multi-site pooling for rare disease studies.

Abstract

Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisition characteristics have been shown to contribute substantially to this predictability. Whether the same holds in brain MRI remains unclear, as anatomical variation and acquisition-dependent contrast are deeply entangled in the image formation process, obscuring the origins of demographic signal. To address this, we propose a controlled framework based on disentangled representation learning, decomposing brain MRI into anatomy-focused representations that suppress acquisition influence and contrast embeddings that capture acquisition-dependent characteristics. Training predictive models for age, sex, and race on full images, anatomical representations, and contrast embeddings allows us to quantify the relative contributions of structure and acquisition to the demographic signal. Across three datasets and multiple MRI sequences, demographic predictability is found to be driven primarily by anatomical variation, with anatomy-focused representations largely preserving the performance of models trained on raw images. Contrast embeddings retain a weaker signal that is dataset-specific and does not generalise across sites. These findings suggest that effective mitigation must explicitly account for the primarily anatomical and secondarily acquisition-dependent origins of demographic signal, ensuring that any bias reduction generalizes robustly across domains.

Sources

Related papers