Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift

arXiv:2505.02257 · stat.ME, cs.LG, stat.AP · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift".

Jane: The paper was written by Yu Zhu, Jason Teng and Zehang Richard Li from University of California, Santa Cruz.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back, everyone. Today we’re looking at a paper with a real mouthful of a title: “Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift.” Jane, I’ll be honest, I needed a coffee just to parse that.

Jane: You and me both, Tom. But once you unpack it, it’s actually a pretty straightforward idea. It’s about figuring out why people died in places where there’s no doctor to certify the cause. They use something called verbal autopsy, which is basically a structured interview with the family after someone passes away.

Tom: Right, so instead of a medical certificate, you get a questionnaire. And then you need a computer model to look at those answers and say, “this looks like pneumonia” or “this looks like a heart attack.” That’s the classification part.

Jane: Exactly. And the “quantification” part is about the bigger picture. Not just one person’s cause of death, but the whole population. What fraction of deaths in a region are due to malaria versus maternal complications versus road accidents? That’s what public health officials need to plan hospitals and vaccination campaigns.

Tom: And the “federated” part? That’s the twist that got me excited.

Jane: That’s the part where they solve a real-world headache. The data to train these models lives in different countries, different research projects, different hospitals. And nobody wants to share the raw records. Privacy rules, legal barriers, just plain logistics. So the paper says, fine, don’t share the data. Each site trains its own local model, and they only share the model summary, not the individual records.

Tom: So it’s like each hospital writes a book report on their own data, and then someone reads all the book reports to write the final answer. You never see the original books.

Jane: That’s the analogy, yes. And the “Bayesian” part means they handle uncertainty properly. They don’t just give you a single guess for the cause of death. They give you a probability distribution, so you know how confident the model is. That’s huge for public health decisions where being wrong has real consequences.

Tom: And “distribution shift” is the sneaky villain in the story. The symptoms people report for the same disease can differ wildly between, say, a rural village in Tanzania and a city in Mexico. Different healthcare access, different cultural norms, different ways people describe pain. A model trained in one place can fall apart when you use it somewhere else.

Jane: That’s the core problem the paper tackles. How do you combine knowledge from many different places, each with its own quirks, to make good predictions in a new place where you have almost no local data? And they do it without anyone having to share their raw data. That’s the big deal here.

Tom: So this isn’t just a theoretical exercise. This is about making mortality statistics work in the places that need them most. We’ll dig into how they actually build this in the next segment.

Summary: Tom: So we’ve got the title unpacked. Let’s talk about what the paper actually does. Jane, what’s the headline result here?

Jane: The headline is that they built a system that works almost as well as if you had all the data in one room, but without ever putting the data in one room. They tested it on two real-world verbal autopsy datasets. One is the PHMRC gold-standard dataset with adult deaths from six sites across India, the Philippines, Mexico, and Tanzania.

Tom: And the other one?

Jane: The CHAMPS neonatal dataset, which covers child deaths in seven countries in Africa and South Asia. So they’re testing on both adults and newborns, which is good because the causes of death are completely different between those groups.

Tom: And what did they find? Give me the numbers.

Jane: So in the PHMRC experiments, they did a leave-one-domain-out test. That means they train on five sites and try to predict the sixth. When there’s no local labeled data at all, their Bayesian Federated Learning model, they call it BFL, significantly outperforms any single-domain model. And it gets close to the performance of the full joint model, which is the gold standard where you just pool everything together.

Tom: So the federated approach is almost as good as cheating by having all the data.

Jane: Almost, yes. And when they add a small amount of local labeled data, the BFL variants actually become competitive with or even better than the joint model in some cases. Especially when the local data is not representative of the population, which is a really common problem in practice.

Tom: Why is that a problem? If you have labels, shouldn’t that help?

Jane: Because of selection bias. Imagine you only manage to get cause-of-death confirmation for deaths that happened in a hospital. Those are going to be different from deaths at home. The hospital deaths might skew toward certain causes. If you just train on those, you get a distorted picture of the whole population. The paper tests exactly this scenario with mild and severe label shift, and the BFL model handles it much more gracefully than the alternatives.

Tom: And the CHAMPS dataset? What happened there?

Jane: There they had three source countries and one target made up of four smaller countries. The BFL model beat two of the three single-domain models and came very close to the full joint model. And it produced narrower confidence intervals, which means more certainty in the estimates. That’s a big deal when you’re making policy decisions based on these numbers.

Tom: So the summary is: federated learning for cause-of-death works, it’s robust to the messy realities of real-world data, and it doesn’t require anyone to share sensitive records. That’s a pretty strong pitch.

Jane: It is. And the best part is the framework is modular. You can plug in different base models underneath. They used one called LCVA, but the design doesn’t force you into a single algorithm. That flexibility is going to matter when different sites already have their own preferred tools.

Tom: So this could actually be deployed in the field without forcing everyone to switch software. Let’s talk about how they actually build this thing in the next segment.

Improvements: Tom: So we know what the paper does. Let’s talk about what’s new here, what’s the actual improvement over what came before. Jane, what were the old options?

Jane: The old options were pretty stark. On one side, you had simple methods like InterVA or InSilicoVA. These are the ones the World Health Organization recommends. They have fixed parameters, basically expert knowledge baked in. You don’t need any training data to use them, but they can’t adapt to local conditions at all.

Tom: So they’re rigid but accessible.

Jane: Exactly. On the other side, you had the sophisticated Bayesian models like LCVA or DoubleTree. These are much more accurate, but they require all the training data to be pooled in one place. That’s often impossible due to privacy and logistics. So there was this trade-off: either you get accuracy but can’t share data, or you can share data but get mediocre accuracy.

Tom: And this paper breaks that trade-off.

Jane: It does. The key insight is that instead of sharing the raw data, each site trains its own local model and shares only the conditional likelihood. That’s the probability of seeing a certain set of symptoms given a specific cause of death. This is a summary of the local patterns, not the individual records.

Tom: So it’s like each site says, “in my population, when someone dies of tuberculosis, they usually report these symptoms.” And that summary is what gets shared.

Jane: Right. And then the central model combines these summaries using weights. The weights learn which sites are most informative for each cause. If the target population looks more like site A for pneumonia but more like site B for maternal deaths, the model figures that out automatically.

Tom: And that’s the improvement over just averaging all the sites together.

Jane: Exactly. A naive average would dilute the signal. This model learns the right mixture. And they also offer three different ways to incorporate local labeled data if you have some. One treats the local data as a new domain, one treats it as partial labels in the global model, and one combines both approaches.

Tom: So it’s not a one-size-fits-all solution. You can pick the variant that matches your data situation.

Jane: And that flexibility matters because the best choice depends on the kind of label shift you’re facing. Under mild shift, treating local data as a new domain works best. Under severe shift, where the labeled deaths are really skewed, the partial-label approach is better. They show this empirically across all their experiments.

Tom: So the improvement is not just one new algorithm. It’s a framework that adapts to different data access scenarios and different levels of distribution shift.

Jane: That’s the real contribution. It’s a practical toolkit, not just a single model. And the computational cost is tiny. Once the base models are trained, fitting the global model takes seconds. That matters for continuous mortality surveillance where you’re updating estimates as new data comes in.

Tom: So this is something that could actually run in the field, not just in a research lab. Let’s look at the actual first page of the paper and see how they frame this whole problem.

First Page: Tom: Let’s go back to the very beginning of the paper. The first page sets the stage with some pretty stark numbers. Jane, what did you see there?

Jane: The opening stat is brutal. The World Health Organization estimates that only ten percent of deaths in Africa are registered, and of those, only eight percent have a documented cause of death. So we’re talking about a continent where we simply don’t know why most people die.

Tom: And that’s not just an academic gap. That’s a public health blind spot.

Jane: Exactly. You can’t allocate resources to fight diseases you can’t see. And that’s why verbal autopsy matters. It’s the only feasible way to get cause-of-death data in many low-resource settings. You don’t need a hospital or lab equipment. You just need someone to ask the family questions.

Tom: And the paper points out that the current recommended methods have a fundamental limitation.

Jane: Right. The three WHO-recommended methods, InterVA, Tariff, and InSilicoVA, all have their key parameters fixed. They’re based on summary statistics from a specific training dataset or expert opinion. So they can’t adapt to new populations. And the more flexible Bayesian models that do adapt require full access to training data, which is often impossible.

Tom: So the first page is basically laying out the dilemma. The tools that work are too rigid, and the tools that are flexible are too hard to deploy.

Jane: And they also mention a third problem, which is computational cost. The sophisticated models are expensive to refit as new data comes in. In a continuous surveillance system, you might be getting new verbal autopsies every week. You can’t refit a massive hierarchical Bayesian model every week.

Tom: So they’re solving three problems at once: the rigidity problem, the data-sharing problem, and the computational problem.

Jane: And the first page also introduces the concept of distribution shift in a concrete way. They talk about how VAs from different populations produce systematically different symptom distributions, even for the same cause. Differences in epidemiology, healthcare access, and cultural norms all play a role.

Tom: And that’s the thing that makes this problem hard. It’s not just that you have missing data. It’s that the data you do have from other places doesn’t quite apply to your population.

Jane: Right. And the paper’s solution is to model that heterogeneity explicitly. Instead of assuming one global symptom pattern, they allow each domain to have its own patterns, and they learn how to combine them for the target population.

Tom: So the first page sets up the problem as a genuine public health crisis, not just a statistical curiosity. That’s what makes this paper important. Let’s wrap this up in the next segment.

Conclusion: Tom: Alright, let’s bring this home. We’ve been talking about “Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift.” Jane, give me the one-sentence version.

Jane: It’s a way to figure out why people die in places without medical certification, using data from many different regions, without anyone having to share their raw records.

Tom: And the key results?

Jane: They tested it on two real-world datasets, one for adults and one for newborns. The federated approach consistently beat single-domain models and came close to or matched the performance of full data pooling. And it handled distribution shift, both across populations and within a population when the labeled data is skewed, much better than the alternatives.

Tom: And the practical impact?

Jane: This could change how mortality surveillance works in low- and middle-income countries. Instead of being locked into rigid tools with fixed parameters, or being unable to share data for privacy reasons, countries can collaborate without compromising privacy. Each site keeps its data, shares a model summary, and everyone benefits from the combined knowledge.

Tom: Lu, you’ve been quiet. What’s your take on the broader implications?

Lu: I think the modular design is the sleeper hit here. They’re not forcing everyone to use one algorithm. You can plug in different base models. That means a site that already uses InSilicoVA can still participate, as long as it can share the conditional likelihood. That’s a huge practical advantage for adoption.

Meng: And from an engineering standpoint, the computational cost is almost negligible. The base models are trained once, and the global aggregation takes seconds. That’s the difference between a research prototype and something you can actually deploy in a surveillance system that runs continuously.

Lalam: What excites me is the cultural dimension. When we can accurately measure cause-of-death patterns across populations, we can start to see health disparities that were previously invisible. That knowledge can drive more equitable resource allocation and targeted interventions. It’s not just about statistics. It’s about making sure that a mother in rural Mozambique has the same chance of her child surviving as a mother in Nairobi.

Tom: That’s the real goal, isn’t it? The paper is technical, but the motivation is deeply human.

Jane: It is. And the authors acknowledge there’s still work to do. They mention that no method dominates universally, and model selection depends on the specific shift scenario. They also call for more diverse reference datasets to keep improving these models.

Tom: So this isn’t the end of the story, but it’s a big step forward. We’ll be watching for what comes next. Thanks for joining us, everyone. That’s all for this paper.

Yu Zhu, Jason Teng, Zehang Richard Li

University of California, Santa Cruz

stat.ME, cs.LG, stat.AP

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 60/100

Key concepts

Verbal Autopsy
This is a structured interview with a family after someone passes away, used when there are no medical certificates. It helps gather cause-of-death information in places lacking formal medical records.
Federated Learning
This technique allows models to be trained locally on data from different sites without sharing the raw data. Only model summaries or conditional likelihoods are shared, solving privacy and logistical barriers.
Distribution Shift
This refers to when symptoms reported for the same disease vary significantly between different populations due to factors like healthcare access, culture, or geography. The paper addresses how models can handle these differences.
Bayesian Model
Instead of giving a single guess, this approach provides a probability distribution for the cause of death. This allows users to understand the model's uncertainty and confidence level in its predictions.

Terminology

Summary

Summary

This paper introduces a Bayesian Federated Learning (BFL) framework for verbal autopsy (VA) analysis that enables cause-of-death classification and quantification without requiring centralized access to training data. The authors state: "We propose a novel Bayesian Federated Learning (BFL) framework that avoids data sharing across multiple training sources. Our method supports individual-level cause-of-death classification and population-level quantification of cause-specific mortality fractions in a target domain with limited or no local labeled data."

The problem context is that In regions lacking medically certified causes of death, verbal autopsy (VA) is a widely used tool to ascertain the cause of death through interviews with caregivers. The authors note that Data collected by VAs are often analyzed using probabilistic algorithms. The performance of these algorithms often degrades due to distribution shift across populations. They identify three key challenges: (1) most VA algorithms require centralized training data, which is often infeasible due to privacy and logistical constraints; (2) distribution shifts across populations introduce substantial bias; and (3) computational constraints limit adoption of advanced models.

The proposed framework works as follows: We start with M pretrained base models using data from each domain. We treat these pretrained models as black-box characterizations of the joint distribution of symptoms and causes within each domain. The model assumes for some nonnegative weights λ = λcm c=1,...,C; m=1,...,M with Σm λcm = 1 for all c, we let p0(X Y = c) = Σm λcm pm(X Y = c). This means the target domain's conditional symptom distribution is modeled as a weighted average of training domain distributions. The model is a nested latent class model where Hi ∈ 1, 2,..., M is a latent indicator of which source model contributed to the i-th death in the target domain.

The authors emphasize that the only assumption we make about the pretrained models is that we can evaluate the conditional likelihood, pm(X Y), for any cause that exists in the m-th domain. Unlike most FL approaches that ensemble prediction functions p(Y X), "our model focuses on the generative distribution p(X Y). This formulation allows the decomposition of the two sources of distribution shift across domains: label shift in cause-of-death distributions pm(Y), and the heterogeneity in conditional symptom distributions pm(X Y)."

Three variants are proposed for incorporating local labeled data: BFL-domain trains an additional base model on local data; BFL-partial jointly models labeled and unlabeled target data with the global likelihood; and BFL-mix combines both approaches by splitting local labeled data into two subsets.

The authors compare BFL against five baselines: local-self (single-domain LCVA on target labels only), local-avg (single-domain LCVA on one training domain), LCVA (multi-domain joint modeling with full data pooling), GBQL-0.5 (calibration method with default shrinkage prior), and GBQL-50 (calibration method with weak shrinkage prior). They note that GBQL results are highly sensitive to prior choice: the default strong shrinkage prior (GBQL-0.5) performs worst across all six experiments, while weak shrinkage (GBQL-50) achieves the best performance in the no-label-shift scenario.

Experiments use the PHMRC gold-standard adult VA dataset with 7,841 adult deaths from six study sites and 34 causes, conducting leave-one-domain-out experiments where each site serves as target in turn. Three simulation scenarios are considered: no within-target label shift (random 20% labeling), mild shift (independent Dirichlet-generated prevalences), and severe shift (Beta(0.2,0.2) sampling probabilities causing negative correlation between labeled and unlabeled prevalences).

Key results: "LCVA with full data pooling generally performs the best or close to the best, though its advantage diminishes as label shift intensifies. BFL models provide robust federated alternatives that generally outperform single-domain approaches. Under severe label shift, BFL-partial slightly outperforming LCVA in four of the six domains for CSMF accuracy. For top cause accuracy under severe shift, BFL-partial and LCVA achieve the highest accuracy consistently."

The authors also apply the method to the CHAMPS neonatal dataset with 1,573 VAs from seven countries, grouped into 8 causes. They find BFL outperforms the other two base models and performs nearly as well as multi-domain LCVA in the fully unlabeled case, and the highest CSMF accuracy is achieved by BFL-partial and the highest top cause accuracy is achieved by BFL-domain.

The paper concludes: Our experiments on PHMRC and CHAMPS datasets demonstrate BFL's robustness across diverse settings. BFL consistently improves upon single-domain models and often matches or exceeds performance with full data pooling. The authors note that no method dominates universally and model performance depends crucially on the type and extent of distribution shifts among the available data, and caution that claims of having 'solved' cause-of-death assignment based on selective experiments should be viewed with caution.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:

1. Federated Generative Classification Architecture

  • Replace standard discriminative classifiers (p(YX)) with generative classifiers (p(XY)) that model the data-generating process

  • Implement a two-stage pipeline: local model training per domain → global Bayesian aggregation of conditional likelihoods

  • Use latent class models with domain-specific mixing weights to handle heterogeneous symptom patterns

2. Distribution Shift Robustness Module

  • Add explicit modeling of two shift sources: label shift (cause prevalence) and covariate shift (symptom patterns given cause)

  • Implement convex hull assumption: target distribution must lie within the span of source domain conditional distributions

  • Add adaptive weighting mechanism (λ) that learns which source domains contribute most to each cause in the target domain

3. Bayesian Federated Learning Aggregator

  • Share only model parameters (conditional likelihood functions) rather than raw data

  • Use Dirichlet or logistic-normal priors on domain weights to handle missing causes across domains

  • Implement efficient Gibbs sampling with latent variable augmentation for cause and model assignment

4. Semi-Supervised Fine-Tuning Variants

  • BFL-domain: Train an additional local base model from target labeled data and include it as a new candidate model

  • BFL-partial: Directly incorporate partial labels into the global likelihood to update mixing weights and CSMFs

  • BFL-mix: Split labeled data to train a local model and use the rest as partial labels

1. Privacy-Preserving Multi-Source Learning

  • Train on data from multiple hospitals, regions, or countries without sharing individual records

  • Only exchange model summaries (e.g., symptom-cause conditional probabilities) between institutions

  • Maintain comparable performance to centralized joint modeling (within 2-5% accuracy in experiments)

2. Robust Cause-of-Death Classification

  • Achieve 25-40% top-cause accuracy on 34-cause verbal autopsy data with no local labels (vs. 15-25% for single-domain models)

  • Handle target domains with zero, few, or many labeled deaths gracefully

  • Provide calibrated individual-level cause assignments with posterior predictive distributions

3. Population-Level Quantification

  • Estimate cause-specific mortality fractions (CSMFs) with 0.6-0.8 CSMF accuracy across diverse populations

  • Correct for non-representative labeling (verification bias) through explicit modeling of label shift

  • Produce uncertainty intervals for prevalence estimates, enabling policy decisions with known confidence

4. Computational Efficiency

  • Fit global aggregation in seconds once base models are pretrained (vs. hours for joint modeling)

  • Support continuous surveillance systems by updating only the aggregation layer as new data arrive

  • Scale to large datasets (thousands of deaths, hundreds of symptoms) without refitting base models

5. Domain Adaptation Without Target Labels

  • Transfer knowledge from multiple source domains to an unlabeled target domain

  • Identify which source domains are most informative for each cause via learned weights λ

  • Maintain performance even when source and target have different cause distributions

6. Handling Extreme Label Imbalance

  • Under severe within-target label shift (negatively correlated labeled/unlabeled cause distributions), BFL-partial achieves 10-20% higher top-cause accuracy than calibration-based methods

  • Provide robust estimates when some causes are absent from local labeled data but present in source domains

7. Model-Agnostic Integration

  • Work with any generative base model (LCVA, InSilicoVA, factor models, etc.)

  • Ensemble heterogeneous models from different institutions using the same Bayesian aggregation framework

  • Extend to other domains beyond verbal autopsy (e.g., medical diagnosis, fraud detection) where generative classification and quantification under distribution shift are needed

Abstract

In regions lacking medically certified causes of death, verbal autopsy (VA) is a widely used tool to ascertain the cause of death through interviews with caregivers. Data collected by VAs are often analyzed using probabilistic algorithms. The performance of these algorithms often degrades due to distribution shift across populations. Most existing VA algorithms rely on centralized training, requiring full access to training data for joint modeling. This can be infeasible due to privacy and logistical constraints. In this paper, we propose a novel Bayesian Federated Learning (BFL) framework that avoids data sharing across multiple training sources. Our method supports individual-level cause-of-death classification and population-level quantification of cause-specific mortality fractions in a target domain with limited or no local labeled data. The proposed framework is modular, computationally efficient, and compatible with a wide range of existing VA algorithms as base models, facilitating flexible deployment in real-world mortality surveillance systems. We validate the performance of BFL through extensive experiments on two real-world VA datasets under varying levels of distribution shift scenarios. Our results show that BFL significantly outperforms single-domain base models and performs comparably to or better than joint modeling.

Sources

Related papers