Federated learning of causal estimands may greatly improve estimation efficiency by leveraging data from multiple study sites, but robustness to heterogeneity and model misspecifications is vital for ensuring validity. We develop a Federated Adaptive Causal Estimation (FACE) framework to incorporate heterogeneous data from multiple sites to provide treatment effect estimation and inference for a flexibly specified target population of interest. FACE accounts for site-level heterogeneity in the distribution of covariates through density ratio weighting. To safely incorporate source sites and avoid negative transfer, we introduce an adaptive weighting procedure via a penalized regression, which achieves both consistency and optimal efficiency. Our strategy is communication-efficient and privacy-preserving, allowing participating sites to share summary statistics only once with other sites. We conduct both theoretical and numerical evaluations of FACE and apply it to conduct a comparative effectiveness study of BNT162b2 (Pfizer) and mRNA-1273 (Moderna) vaccines on COVID-19 outcomes in U.S. veterans using electronic health records from five VA regional sites. We show that compared to traditional methods, FACE meaningfully increases the precision of treatment effect estimates, with reductions in standard errors ranging from $26\%$ to $67\%$.
翻译:联邦学习因果估计量可通过整合多个研究站点的数据显著提升估计效率,但确保对异质性和模型误设定的鲁棒性至关重要。我们开发了联邦自适应因果估计(FACE)框架,用于整合来自多个站点的异质性数据,从而对灵活指定的目标人群进行治疗效果估计和推断。FACE通过密度比加权方法处理站点间协变量分布的异质性。为安全利用源站点并避免负迁移,我们引入基于惩罚回归的自适应加权程序,该程序同时实现了估计一致性和最优效率。该策略具有通信高效性和隐私保护性,允许参与站点仅需与其他站点共享一次汇总统计量。我们对FACE进行了理论与数值评估,并将其应用于一项比较效果研究:利用五个VA区域站点的电子健康记录,评估BNT162b2(辉瑞)与mRNA-1273(莫德纳)疫苗对美国退伍军人COVID-19结局的影响。结果表明,与传统方法相比,FACE显著提高了治疗效果估计的精度,标准误降低幅度达26%至67%。