A novel mixture cure frailty model is introduced for handling censored survival data. Mixture cure models are preferable when the existence of a cured fraction among patients can be assumed. However, such models are heavily underexplored: frailty structures within cure models remain largely undeveloped, and furthermore, most existing methods do not work for high-dimensional datasets, when the number of predictors is significantly larger than the number of observations. In this study, we introduce a novel extension of the Weibull mixture cure model that incorporates a frailty component, employed to model an underlying latent population heterogeneity with respect to the outcome risk. Additionally, high-dimensional covariates are integrated into both the cure rate and survival part of the model, providing a comprehensive approach to employ the model in the context of high-dimensional omics data. We also perform variable selection via an adaptive elastic-net penalization, and propose a novel approach to inference using the expectation-maximization (EM) algorithm. Extensive simulation studies are conducted across various scenarios to demonstrate the performance of the model, and results indicate that our proposed method outperforms competitor models. We apply the novel approach to analyze RNAseq gene expression data from bulk breast cancer patients included in The Cancer Genome Atlas (TCGA) database. A set of prognostic biomarkers is then derived from selected genes, and subsequently validated via both functional enrichment analysis and comparison to the existing biological literature. Finally, a prognostic risk score index based on the identified biomarkers is proposed and validated by exploring the patients' survival.
翻译:本文提出了一种新颖的混合治愈脆弱模型,用于处理删失生存数据。当可以假设患者中存在治愈比例时,混合治愈模型更为适用。然而,这类模型的研究尚严重不足:治愈模型中的脆弱结构仍远未发展完善,且大多数现有方法无法处理预测变量数量远大于观测数量的高维数据集。本研究引入了一种威布尔混合治愈模型的新扩展,该模型纳入脆弱成分以模拟潜在的人群异质性对结局风险的影响。此外,高维协变量被整合到模型的治愈率和生存两部分中,为在高维组学数据背景下应用该模型提供了全面方法。我们还通过自适应弹性网惩罚进行变量选择,并基于期望最大化(EM)算法提出了一种新的推断方法。通过多种场景下的广泛模拟研究验证了模型性能,结果表明所提方法优于竞争模型。我们将该新方法应用于分析癌症基因组图谱(TCGA)数据库中批量乳腺癌患者的RNAseq基因表达数据,从所选基因中推导出一组预后生物标志物,并进一步通过功能富集分析和与现有生物学文献的比对进行验证。最后,基于所识别生物标志物提出了预后风险评分指数,并通过探索患者生存情况加以验证。