Shapley Value Estimation for Multi-Site Data with Blockwise-Missing Features
Siqi Li, Wangxuan Fan, Yiming Li, Doudou Zhou, Molei Liu
Abstract
Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley estimators generally assume that observations used to evaluate the coalitional game are fully observed under a common feature space. This assumption is routinely violated in multi-site studies across biomedicine, social science, and environmental monitoring, where institutions record different features under different protocols, producing systematic blockwise missingness across sources. We first show that the standard remedy of imputing missing features before computing Shapley values introduces systematic, coalition-dependent bias into the resulting attributions. We then propose FUSHAP (Fusion Shapley Attribution from Partially-observed data), a method that leverages partially-observed auxiliary sites to reduce the variance of a preliminary single-site Shapley estimate without imputation. A permutation-based screening step detects and excludes sites whose data distributions are incompatible with the target population. In synthetic experiments, FUSHAP achieves 3--8× lower MSE than the single-site estimator and 2--3× lower MSE than imputation baselines without incurring imputation-induced bias, and the screening procedure identifies misaligned sites with 82\% power at moderate misalignment and 100\% for strong misalignment. On multi-site air quality and multi-center clinical data, FUSHAP reduces MSE by approximately 3--7× relative to the single-site estimator; in the clinical application, standard imputation can increase MSE above the single-site baseline.
Create a lesson
Related papers
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
Sho Kawano, Zehang Richard Li, Paul A. Parker
TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models
Jingbo Liu, Zhiyuan Yu
Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs
Zhenlin Yao, Wei Xiong
Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning
Weiwei Wang, Yuqiang Li, Xianyi Wu et al.
Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks
Xianjun Li, Yunfei Yang
Next-token functional estimation
Milind Nakul, Vidya Muthukumar, Ashwin Pananjady