The Cox model is an indispensable tool for time-to-event analysis, particularly in biomedical research. However, medicine is undergoing a profound transformation, generating data at an unprecedented scale, which opens new frontiers to study and understand diseases. With the wealth of data collected, new challenges for statistical inference arise, as datasets are often high dimensional, exhibit an increasing number of measurements at irregularly spaced time points, and are simply too large to fit in memory. Many current implementations for time-to-event analysis are ill-suited for these problems as inference is computationally demanding and requires access to the full data at once. Here we propose a Bayesian version for the counting process representation of Cox's partial likelihood for efficient inference on large-scale datasets with millions of data points and thousands of time-dependent covariates. Through the combination of stochastic variational inference and a reweighting of the log-likelihood, we obtain an approximation for the posterior distribution that factorizes over subsamples of the data, enabling the analysis in big data settings. Crucially, the method produces viable uncertainty estimates for large-scale and high-dimensional datasets. We show the utility of our method through a simulation study and an application to myocardial infarction in the UK Biobank.
翻译:Cox模型是时间-事件分析中不可或缺的工具,尤其在生物医学研究中。然而,医学领域正经历深刻变革,产生的数据规模前所未有,这为研究疾病开辟了新前沿。随着海量数据的收集,统计推断面临新挑战:数据集往往呈现高维性、不规则时间点测量值激增,且规模过大无法一次性载入内存。当前多数时间-事件分析实现方法因推断计算量庞大且需一次性访问完整数据,难以应对此类问题。本文提出一种基于计数过程表示的Cox部分似然贝叶斯方法,可高效推断包含百万级数据点和数千个时变协变量的大规模数据集。通过结合随机变分推断与对数似然重加权,我们获得因子化分解于数据子样本的后验分布近似,从而支持大数据场景分析。关键在于,该方法能为大规模高维数据集提供可靠的估计不确定性。通过模拟研究与英国生物银行心肌梗死数据的应用,我们验证了该方法的实用性。