Accurate predictions, as with machine learning, may not suffice to provide optimal healthcare for every patient. Indeed, prediction can be driven by shortcuts in the data, such as racial biases. Causal thinking is needed for data-driven decisions. Here, we give an introduction to the key elements, focusing on routinely-collected data, electronic health records (EHRs) and claims data. Using such data to assess the value of an intervention requires care: temporal dependencies and existing practices easily confound the causal effect. We present a step-by-step framework to help build valid decision making from real-life patient records by emulating a randomized trial before individualizing decisions, eg with machine learning. Our framework highlights the most important pitfalls and considerations in analysing EHRs or claims data to draw causal conclusions. We illustrate the various choices in studying the effect of albumin on sepsis mortality in the Medical Information Mart for Intensive Care database (MIMIC-IV). We study the impact of various choices at every step, from feature extraction to causal-estimator selection. In a tutorial spirit, the code and the data are openly available.
翻译:基于机器学习的精确预测可能不足以实现每位患者的最佳诊疗,因为预测结果可能受到数据中的捷径偏差(如种族偏见)影响。数据驱动的决策需要因果思维。本文针对常规收集的电子健康记录(EHR)和理赔数据,系统介绍了因果推理的核心要素。利用此类数据评估干预措施的价值需谨慎:时间依赖性和既有实践容易混淆因果效应。我们提出一个分步框架,通过模拟随机试验来构建基于真实患者记录的合理决策,进而实现个体化干预(例如借助机器学习)。该框架重点阐述了分析EHR或理赔数据时推导因果结论的关键陷阱与考量。我们以重症监护医学信息数据库(MIMIC-IV)中白蛋白对脓毒症死亡率的影响研究为例,展示了从特征提取到因果估计器选择各环节不同决策的影响。秉承教学宗旨,本文代码与数据均公开可用。