Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
summary
In short
The episode discusses a paper titled "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Researchers found that demographic predictability in brain MRI is primarily driven by anatomical structure, not scanner contrast. The anatomy-focused representations maintained most predictive power across datasets, while contrast embeddings showed site-specific fragility.
Key concepts
- Anatomy Representation
- This representation captures the actual shape and structure of the brain from an MRI scan. The study found that this component preserved almost all predictive power for demographic factors like sex and age, suggesting structural features are the main source of demographic signal.
- Contrast Embedding
- This embedding captures information about the way the image was taken, such as scanner settings and protocol. The research showed that contrast embeddings were weaker predictors of demographics and were fragile when tested across different datasets, indicating this signal is site-specific.
- Disentangling Models (MR-CLIP and DIST-CLIP)
- These models split a single brain scan into two separate representations: one capturing anatomy and another capturing contrast. This method allowed researchers to test which component—anatomy or contrast—was responsible for demographic predictions, revealing the source of the signal.
- Cross-Dataset Generalization
- This refers to testing if a model trained on one dataset (e.g., HCP) performs well on a different dataset (e.g., OASIS). The study showed that while contrast signals were site-specific and failed to generalize, the anatomy signal transferred more robustly between sites.
Terminology used across episodes
This episode discusses
- Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast · Paper Radio
- Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
- Metadata-Aligned 3D MRI Representations for Contrast Understanding and Quality Control
- Deep Residual Learning for Image Recognition
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settings
The paper
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast · Read on arXiv
Mehmet Yigit Avci, Akshit Achara, Andrew King, Jorge Cardoso
King's College London
Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisition characteristics have been shown to contribute substantially to this predictability. Whether the same holds in brain MRI remains unclear, as anatomical variation and acquisition-dependent contrast are deeply entangled in the image formation process, obscuring the origins of demographic signal. To address this, we propose a controlled framework based on disentangled representation learning, decomposing brain MRI into anatomy-focused representations that suppress acquisition influence and contrast embeddings that capture acquisition-dependent characteristics. Training predictive models for age, sex, and race on full images, anatomical representations, and contrast embeddings allows us to quantify the relative contributions of structure and acquisition to the demographic signal. Across three datasets and multiple MRI sequences, demographic predictability is found to be driven primarily by anatomical variation, with anatomy-focused representations largely preserving the performance of models trained on raw images. Contrast embeddings retain a weaker signal that is dataset-specific and does not generalise across sites. These findings suggest that effective mitigation must explicitly account for the primarily anatomical and secondarily acquisition-dependent origins of demographic signal, ensuring that any bias reduction generalizes robustly across domains.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast".
Jane: The paper was written by Mehmet Yigit Avci, Akshit Achara, Andrew King and Jorge Cardoso from King's College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: Welcome back, everyone. Tom here with Jane, and we are kicking off our discussion of a brand new paper on arXiv. It’s called "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, I have to say, the title alone is a mouthful, but the question it asks is huge.
Jane: It really is, Tom. And the question is basically this: when a computer looks at a brain MRI and can guess your age, your sex, even your race, where is that information actually coming from? Is it in the physical structure of your brain, or is it hiding in the way the picture was taken?
Tom: Right, and that matters because these models are being used for real clinical tasks like diagnosing Alzheimer's. If they're picking up on demographic clues instead of actual disease, that's a bias problem. But nobody had really pinned down the source of that signal in brain MRI before.
Jane: Exactly. In chest X-rays, researchers found that a lot of the demographic signal comes from the acquisition settings—the machine, the protocol, the exposure. The team here, led by Mehmet Yigit Avci and Akshit Achara at King's College London, wanted to see if the same was true for brain MRI.
Tom: And they built a clever way to test it. They used two models, MR-CLIP and DIST-CLIP, which can take a single brain scan and split it into two separate representations. One captures the anatomy, the actual shape and structure of the brain. The other captures the contrast, which is basically the imaging fingerprint of the scanner and protocol.
Jane: So it's like separating the recipe from the chef. The anatomy is the ingredients, the contrast is the cooking style. And then they trained predictors for age, sex, and race on each of those components separately, plus on the raw full image.
Tom: And the results were pretty striking. Across three datasets—OASIS, ADNI, and HCP—the anatomy representations preserved almost all of the predictive power. For sex prediction, raw images hit about ninety-two to ninety-five percent balanced accuracy, and the anatomy-only representations were right there at ninety-three to ninety-six percent.
Jane: Meanwhile, the contrast embeddings were much weaker. They still had some signal, but it was clearly secondary. So the big takeaway here is that in brain MRI, the demographic signal is primarily anatomical, not a scanner artifact.
Tom: That's a big deal because it flips the script from the chest X-ray findings. It means you can't just harmonize the images or normalize the intensities and expect the bias to disappear. The information is baked into the structure itself.
Jane: And that has real consequences for how we build fair clinical models. But before we get into the fixes, I want to dig into how they actually proved this disentanglement works. That's coming up next.
Tom: Stay with us.
Paper discussion segment 2: Jane: Welcome back. We're still on "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Tom, we left off saying the anatomy carries most of the signal. But how do we know the models actually separated anatomy from contrast cleanly?
Tom: Great question, and the paper has a really neat validation. They took five hundred subjects from OASIS and used SynthSeg to get ground-truth brain measurements like intracranial volume and ventricle volume. Then they probed each representation to see what it actually encoded.
Jane: And the results were stark. The anatomy representation predicted intracranial volume with an R-squared of zero point nine five three, almost perfect. But when they tried to predict which scanner or protocol the image came from, it only got forty-two point five percent balanced accuracy, where chance was six point seven percent across fifteen clusters. So it really is anatomy-focused.
Tom: On the flip side, the contrast embedding was much better at guessing the acquisition protocol, hitting seventy point two percent, but it was noticeably worse at anatomy, with an R-squared of zero point eight zero five for intracranial volume. So the separation is real, though not perfect.
Jane: And here's a subtle but important point. The metadata-only embedding, which comes from DICOM parameters and never sees the image, already predicted intracranial volume with an R-squared of zero point seven three nine. That tells us a lot of the residual anatomical signal in the contrast embedding isn't actually image leakage—it's just that certain scanners tend to be used on certain populations.
Tom: That's a really insightful observation. It means some of the demographic signal we see in contrast is just a proxy for site-specific demographics. You scan a wealthier, healthier population at one site, and the scanner settings correlate with that population's brain characteristics.
Jane: Exactly. So when they trained the full models, the raw images performed best, as you'd expect. Age MAE was around two point six seven years in HCP, sex balanced accuracy up to zero point nine six. But the anatomy representations were nearly identical. And the contrast embeddings, while weaker, still had non-trivial signal—sex at zero point seven one in OASIS, for example.
Tom: So within a single dataset, contrast does carry some demographic information. But the real test is generalization. If you train on one site and test on another, does that contrast signal hold up?
Jane: And that's where the story gets even more interesting. We'll get into those cross-dataset results right after this.
Paper discussion segment 3: Tom: Back with "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, we just said within a dataset, contrast has some signal. But the cross-dataset results really tell the story.
Jane: They do. When they trained on ADNI and tested on OASIS, sex prediction from raw images got seventy-six percent balanced accuracy. The anatomy representation actually improved that to eighty-four percent. But the contrast embedding collapsed to seventy percent, and for race it dropped to fifty percent, which is pure chance.
Tom: That pattern repeats across nearly every transfer pair. The contrast signal is dataset-specific. It doesn't generalize. The anatomy signal, on the other hand, transfers much more robustly. For example, training on HCP and testing on OASIS, anatomy sex prediction hit seventy-one percent while raw was seventy-two percent—basically the same.
Jane: And age prediction under cross-dataset shift was brutal. Training on HCP and testing on ADNI gave a mean absolute error of forty-two years. That's essentially random. But even there, the anatomy representation didn't do worse than raw—it was forty-four years, so comparable.
Tom: Right, so the anatomy isn't a magic bullet for domain shift, but it doesn't add extra fragility either. The key finding is that the acquisition-driven signal is fragile and site-specific, while the structural signal is more stable.
Jane: They also looked at multiple sequences within OASIS—T1w, T2w, and FLAIR. And the same pattern held. Anatomy representations stayed close to raw performance across all sequences. For FLAIR, anatomy sex prediction actually beat raw, eighty-three percent versus seventy-one percent.
Tom: That's a fascinating result. It suggests that when the raw image is noisy or the contrast is less informative, the anatomy representation can be more robust because it strips away the irrelevant acquisition noise.
Jane: So what does this mean for bias mitigation? The authors are pretty clear. If you want to reduce demographic predictability in brain MRI models, you can't just focus on harmonizing intensities or normalizing contrast. The signal is in the structure.
Tom: And that's a harder problem. You need to identify which specific anatomical features are driving the demographic signal and decide whether they're clinically relevant or not. For Alzheimer's, some structural changes are exactly what you want to detect. So you can't just throw away anatomy.
Jane: Right, it's a surgical approach, not a blunt one. And the authors suggest that acquisition-related bias should still be addressed, but mainly when it introduces additional demographic information in site-specific settings.
Tom: Before we wrap up, I want to bring in Lu and Meng to get their take on the practical side. Lu, you've been listening—what excites you about this?
Lu: I think the framework is the real contribution here. It gives us a way to decompose the signal and ask mechanistic questions. That's rare in this field. Most bias papers just say "there's bias," but this one says "here's where it lives and here's why."
Meng: And from an engineering standpoint, the fact that the anatomy representation is robust across sites is promising. It means we could potentially use it as a preprocessing step to reduce spurious correlations without losing task-relevant information. But we'd need to validate that on actual clinical outcomes, not just demographic prediction.
Jane: That's a perfect segue to our conclusion. Let's wrap this up.
Conclusion: Tom: Alright, we're closing out our discussion of "Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast." Jane, give us the final summary.
Jane: The paper asked a simple but profound question: when models predict demographics from brain MRI, where does the signal come from? And the answer, across three datasets and multiple sequences, is that it's primarily anatomical. The anatomy-focused representations preserved nearly all the predictive power of raw images, while contrast embeddings were weaker and failed to generalize across sites.
Tom: And that's a direct contrast to findings in chest X-rays, where acquisition played a bigger role. So this paper really advances our understanding of modality-specific bias sources.
Jane: The implications are clear. Bias mitigation in brain MRI can't rely on intensity harmonization alone. It has to engage with the anatomical features themselves, and it has to be careful not to throw away clinically meaningful structure.
Tom: There are limitations, of course. The datasets are imbalanced, especially ADNI with only five point two percent Black participants. Race prediction results should be taken with a grain of salt. But the framework is solid and the methodology is rigorous.
Meng: I'd add that the cross-dataset results are the most actionable part. They show that contrast-based demographic signal is a site artifact, which means site-specific mitigation strategies could work for acquisition bias, but anatomical bias needs a different approach.
Lu: And I'd say the next step is to connect this to downstream clinical tasks. Does reducing demographic predictability in the representation actually reduce disparities in Alzheimer's detection? That's the million-dollar question.
Jane: Exactly. This paper gives us the map, but we still need to navigate the territory. We'll be watching for follow-up work.
Tom: Thanks for joining us, everyone. We're saying goodbye to this paper and getting ready to dive into the next one. Stay curious.
Jane: See you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization