Survival analysis serves as a fundamental component in numerous healthcare applications, where the determination of the time to specific events (such as the onset of a certain disease or death) for patients is crucial for clinical decision-making. Scoring systems are widely used for swift and efficient risk prediction. However, existing methods for constructing survival scores presume that data originates from a single source, posing privacy challenges in collaborations with multiple data owners. We propose a novel framework for building federated scoring systems for multi-site survival outcomes, ensuring both privacy and communication efficiency. We applied our approach to sites with heterogeneous survival data originating from emergency departments in Singapore and the United States. Additionally, we independently developed local scores at each site. In testing datasets from each participant site, our proposed federated scoring system consistently outperformed all local models, evidenced by higher integrated area under the receiver operating characteristic curve (iAUC) values, with a maximum improvement of 11.6%. Additionally, the federated score's time-dependent AUC(t) values showed advantages over local scores, exhibiting narrower confidence intervals (CIs) across most time points. The model developed through our proposed method exhibits effective performance on each local site, signifying noteworthy implications for healthcare research. Sites participating in our proposed federated scoring model training gained benefits by acquiring survival models with enhanced prediction accuracy and efficiency. This study demonstrates the effectiveness of our privacy-preserving federated survival score generation framework and its applicability to real-world heterogeneous survival data.
翻译:生存分析是众多医疗应用的基础组成部分,其中确定患者发生特定事件(如某种疾病发作或死亡)的时间对于临床决策至关重要。评分系统被广泛用于快速高效的风险预测。然而,现有构建生存评分的方法假设数据来源于单一来源,这在多数据方合作中带来了隐私挑战。我们提出了一种新颖的框架,用于构建多站点生存结局的联邦评分系统,同时确保隐私和通信效率。我们将该方法应用于来自新加坡和美国急诊科的异质性生存数据站点。此外,我们独立地在每个站点开发了本地评分。在来自每个参与站点的测试数据集中,我们提出的联邦评分系统始终优于所有本地模型,这通过更高的受试者工作特征曲线下综合面积(iAUC)值得以证明,最大提升了11.6%。此外,联邦评分的时间依赖性AUC(t)值显示出优于本地评分的优势,在大多数时间点上表现出更窄的置信区间(CIs)。通过我们提出方法开发的模型在每个本地站点上都表现出有效性能,这对医疗研究具有重要意义。参与我们提出的联邦评分模型训练的站点通过获得具有更高预测精度和效率的生存模型而获益。本研究证明了我们隐私保护的联邦生存评分生成框架的有效性及其对真实世界异质性生存数据的适用性。