Surrogate variables in electronic health records (EHR) play an important role in biomedical studies due to the scarcity or absence of chart-reviewed gold standard labels, under which supervised methods only using labeled data poorly perform poorly. Meanwhile, synthesizing multi-site EHR data is crucial for powerful and generalizable statistical learning but encounters the privacy constraint that individual-level data is not allowed to be transferred from the local sites, known as DataSHIELD. In this paper, we develop a novel approach named SASH for Surrogate-Assisted and data-Shielding High-dimensional integrative regression. SASH leverages sizable unlabeled data with EHR surrogates predictive of the response from multiple local sites to assist the training with labeled data and largely improve statistical efficiency. It first extracts a preliminary supervised estimator to realize convex training of a regularized single index model for the surrogate at each local site and then aggregates the fitted local models for accurate learning of the target outcome model. It protects individual-level information from the local sites through summary-statistics-based data aggregation. We show that under mild conditions, our method attains substantially lower estimation error rates than the supervised or local semi-supervised methods, as well as the asymptotic equivalence to the ideal individual patient data pooled estimator (IPD) only available in the absence of privacy constraints. Through simulation studies, we demonstrate that SASH outperforms all existing supervised or SS federated approaches and performs closely to IPD. Finally, we apply our method to develop a high dimensional genetic risk model for type II diabetes using large-scale biobank data sets from UK Biobank and Mass General Brigham, where only a small fraction of subjects from the latter has been labeled via chart reviewing.
翻译:摘要:电子健康记录(EHR)中的代理变量在生物医学研究中扮演重要角色,因为经过图表审核的黄金标准标签稀缺或缺失,此时仅使用标注数据的监督方法性能较差。同时,整合多中心EHR数据对于构建强大且泛化能力强的统计学习模型至关重要,但面临隐私约束——不允许从本地站点传输个体级数据,这被称为DataSHIELD。在本文中,我们提出了一种名为SASH(代理辅助与数据屏蔽高维积分回归)的新方法。SASH利用来自多个本地站点的、可预测响应的EHR代理变量的大规模未标注数据,辅助标注数据的训练,并显著提升统计效率。该方法首先提取初步监督估计量,以实现每个本地站点代理变量的正则化单指标模型的凸训练,然后聚合拟合后的本地模型,以精确学习目标结果模型。通过基于汇总统计的数据聚合,保护本地站点的个体级信息。我们证明,在温和条件下,本方法的估计误差率显著低于监督方法或本地半监督方法,并且在渐近意义上等价于仅在无隐私约束条件下可用的理想个体患者数据池化估计器(IPD)。通过模拟研究,我们展示SASH优于所有现有的监督或半监督联邦学习方法,且性能接近IPD。最后,我们将该方法应用于使用英国生物银行和马萨诸塞州总医院布里格姆的大规模生物银行数据集构建II型糖尿病的高维遗传风险模型,其中仅后者的一小部分受试者通过图表审核进行了标注。