Over the last decades, many prognostic models based on artificial intelligence techniques have been used to provide detailed predictions in healthcare. Unfortunately, the real-world observational data used to train and validate these models are almost always affected by biases that can strongly impact the outcomes validity: two examples are values missing not-at-random and selection bias. Addressing them is a key element in achieving transportability and in studying the causal relationships that are critical in clinical decision making, going beyond simpler statistical approaches based on probabilistic association. In this context, we propose a novel approach that combines selection diagrams, missingness graphs, causal discovery and prior knowledge into a single graphical model to estimate the cardiovascular risk of adolescent and young females who survived breast cancer. We learn this model from data comprising two different cohorts of patients. The resulting causal network model is validated by expert clinicians in terms of risk assessment, accuracy and explainability, and provides a prognostic model that outperforms competing machine learning methods.
翻译:过去几十年来,基于人工智能技术的多种预后模型已被用于提供医疗领域的详细预测。然而,用于训练和验证这些模型的实际观察性数据几乎总是受到偏差的影响,这些偏差可能严重影响结果的效度:两个例子是非随机缺失值和选择性偏差。解决这些问题对于实现可迁移性以及研究临床决策中至关重要的因果关系具有关键意义,这超越了基于概率关联的简单统计方法。在此背景下,我们提出了一种新颖的方法,将选择图、缺失图、因果发现和先验知识整合到一个统一的图模型中,用于评估患有乳腺癌后幸存的青少年及年轻女性的心血管风险。我们根据包含两个不同患者队列的数据学习该模型。最终得到的因果网络模型由临床专家在风险评估、准确性和可解释性方面进行验证,并提供了一种优于竞争性机器学习方法的预后模型。