Over the last decades, many prognostic models based on artificial intelligence techniques have been used to provide detailed predictions in healthcare. Unfortunately, the real-world observational data used to train and validate these models are almost always affected by biases that can strongly impact the outcomes validity: two examples are values missing not-at-random and selection bias. Addressing them is a key element in achieving transportability and in studying the causal relationships that are critical in clinical decision making, going beyond simpler statistical approaches based on probabilistic association. In this context, we propose a novel approach that combines selection diagrams, missingness graphs, causal discovery and prior knowledge into a single graphical model to estimate the cardiovascular risk of adolescent and young females who survived breast cancer. We learn this model from data comprising two different cohorts of patients. The resulting causal network model is validated by expert clinicians in terms of risk assessment, accuracy and explainability, and provides a prognostic model that outperforms competing machine learning methods.
翻译:过去几十年,基于人工智能技术的多种预后模型已广泛应用于医疗领域,提供详细的预测结果。然而,用于训练和验证这些模型的实际观测数据几乎总会受到偏差影响,严重削弱结果的有效性:两个典型例子是非随机缺失值和选择偏差。解决这些问题,是实现模型可迁移性和研究临床决策中关键的因果关系的重要环节,这超越了基于概率关联的简单统计方法。在此背景下,我们提出了一种新颖方法,将选择图、缺失图、因果发现和先验知识整合到一个统一的图模型中,用于评估罹患乳腺癌后存活青少年及年轻女性的心血管风险。我们从包含两个不同患者队列的数据中学习该模型。最终得到的因果网络模型由临床专家在风险评估、准确性和可解释性方面进行验证,并展现出优于其他竞争机器学习方法的预后模型性能。