Censoring is the central problem in survival analysis where either the time-to-event (for instance, death), or the time-tocensoring (such as loss of follow-up) is observed for each sample. The majority of existing machine learning-based survival analysis methods assume that survival is conditionally independent of censoring given a set of covariates; an assumption that cannot be verified since only marginal distributions is available from the data. The existence of dependent censoring, along with the inherent bias in current estimators has been demonstrated in a variety of applications, accentuating the need for a more nuanced approach. However, existing methods that adjust for dependent censoring require practitioners to specify the ground truth copula. This requirement poses a significant challenge for practical applications, as model misspecification can lead to substantial bias. In this work, we propose a flexible deep learning-based survival analysis method that simultaneously accommodate for dependent censoring and eliminates the requirement for specifying the ground truth copula. We theoretically prove the identifiability of our model under a broad family of copulas and survival distributions. Experiments results from a wide range of datasets demonstrate that our approach successfully discerns the underlying dependency structure and significantly reduces survival estimation bias when compared to existing methods.
翻译:删失是生存分析中的核心问题,即每个样本只能观测到事件发生时间(如死亡)或删失时间(如失访)。现有基于机器学习的生存分析方法大多假设在给定协变量条件下生存时间与删失时间条件独立——该假设无法验证,因为数据仅提供边际分布。各类应用已证实相依删失的存在及其导致当前估计器的固有偏差,凸显了采用更精细方法的必要性。然而,现有调整相依删失的方法要求研究者指定真实Copula函数。这一要求给实际应用带来重大挑战,因为模型误设可能导致显著偏差。本文提出一种基于深度学习的灵活生存分析方法,既能同时适应相依删失场景,又无需指定真实Copula函数。我们从理论上证明,在宽泛的Copula族和生存分布条件下,该模型具有可识别性。多组数据集的实验结果表明,与现有方法相比,本方法能成功识别潜在依赖结构,并显著降低生存估计偏差。