In recent years, the development of technologies for causal inference with privacy preservation of distributed data has gained considerable attention. Many existing methods for distributed data focus on resolving the lack of subjects (samples) and can only reduce random errors in estimating treatment effects. In this study, we propose a data collaboration quasi-experiment (DC-QE) that resolves the lack of both subjects and covariates, reducing random errors and biases in the estimation. Our method involves constructing dimensionality-reduced intermediate representations from private data from local parties, sharing intermediate representations instead of private data for privacy preservation, estimating propensity scores from the shared intermediate representations, and finally, estimating the treatment effects from propensity scores. Through numerical experiments on both artificial and real-world data, we confirm that our method leads to better estimation results than individual analyses. While dimensionality reduction loses some information in the private data and causes performance degradation, we observe that sharing intermediate representations with many parties to resolve the lack of subjects and covariates sufficiently improves performance to overcome the degradation caused by dimensionality reduction. Although external validity is not necessarily guaranteed, our results suggest that DC-QE is a promising method. With the widespread use of our method, intermediate representations can be published as open data to help researchers find causalities and accumulate a knowledge base.
翻译:近年来,面向分布式数据的隐私保护因果推断技术发展备受关注。现有多数分布式方法仅着眼于解决样本量不足问题,仅能降低治疗效果估计中的随机误差。本研究提出数据协作准实验(DC-QE)方法,可同时解决样本与协变量双重缺失问题,从而降低估计中的随机误差与偏差。该方法包括:从各参与方的私有数据中构建降维中间表示;通过共享中间表示而非原始数据实现隐私保护;基于共享的中间表示估计倾向得分;最终利用倾向得分估计治疗效果。在人工数据集与真实数据上的数值实验表明,该方法相比独立分析可获得更优的估计结果。虽然降维过程会损失私有数据中的部分信息并导致性能下降,但我们观察到,通过多方共享中间表示以弥补样本与协变量的缺失,可充分提升性能以抵消降维带来的负面影响。尽管外部效度未必有保障,但研究结果表明DC-QE是一种具有潜力的方法。随着该方法的广泛应用,中间表示可作为开放数据发布,助力研究者发现因果关系并积累知识库。