Shapley Value Estimation for Multi-Site Data with Blockwise-Missing Features
stat.ML, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/siqili0325/FUSHAP
License: http://creativecommons.org/licenses/by/4.0/
The gist: Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley estimators generally assume that observations used to
Terminology
Abstract
Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley estimators generally assume that observations used to evaluate the coalitional game are fully observed under a common feature space. This assumption is routinely violated in multi-site studies across biomedicine, social science, and environmental monitoring, where institutions record different features under different protocols, producing systematic blockwise missingness across sources. We first show that the standard remedy of imputing missing features before computing Shapley values introduces systematic, coalition-dependent bias into the resulting attributions. We then propose FUSHAP (Fusion Shapley Attribution from Partially-observed data), a method that leverages partially-observed auxiliary sites to reduce the variance of a preliminary single-site Shapley estimate without imputation. A permutation-based screening step detects and excludes sites whose data distributions are incompatible with the target population. In synthetic experiments, FUSHAP achieves 3 -- 8 times lower MSE than the single-site estimator and 2 -- 3 times lower MSE than imputation baselines without incurring imputation-induced bias, and the screening procedure identifies misaligned sites with 82% power at moderate misalignment and 100% for strong misalignment. On multi-site air quality and multi-center clinical data, FUSHAP reduces MSE by approximately 3 -- 7 times relative to the single-site estimator; in the clinical application, standard imputation can increase MSE above the single-site baseline.
Sources
- SIM-Shapley: A Stable and Computationally Efficient Approach to Shapley Value Approximation
- Efficient Semiparametric Inference for Distributed Data with Blockwise Missingness
- Modular Regression: Improving Linear Models by Incorporating Auxiliary Data
- Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction
- Adaptive and Efficient Learning with Blockwise Missing and Semi-Supervised Data
- What Is a Good Imputation Under MAR Missingness?
- Distribution Shift in Missing Data Imputation: A Risk-Based Perspective and Importance-Weighted Correction under MAR
- Explainability of Machine Learning Models under Missing Data
- A Principled Approach to Data Valuation for Federated Learning
- Explaining medical AI performance disparities across sites with confounder Shapley value analysis
- Towards Efficient Inference under Nonmonotone Missingness with General Imputation
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey