Performance monitoring of machine learning (ML)-based risk prediction models in healthcare is complicated by the issue of confounding medical interventions (CMI): when an algorithm predicts a patient to be at high risk for an adverse event, clinicians are more likely to administer prophylactic treatment and alter the very target that the algorithm aims to predict. A simple approach is to ignore CMI and monitor only the untreated patients, whose outcomes remain unaltered. In general, ignoring CMI may inflate Type I error because (i) untreated patients disproportionally represent those with low predicted risk and (ii) evolution in both the model and clinician trust in the model can induce complex dependencies that violate standard assumptions. Nevertheless, we show that valid inference is still possible if one monitors conditional performance and if either conditional exchangeability or time-constant selection bias hold. Specifically, we develop a new score-based cumulative sum (CUSUM) monitoring procedure with dynamic control limits. Through simulations, we demonstrate the benefits of combining model updating with monitoring and investigate how over-trust in a prediction model may delay detection of performance deterioration. Finally, we illustrate how these monitoring methods can be used to detect calibration decay of an ML-based risk calculator for postoperative nausea and vomiting during the COVID-19 pandemic.
翻译:基于机器学习(ML)的风险预测模型在医疗领域的性能监测因混杂医疗干预(CMI)问题而变得复杂:当算法预测患者面临不良事件高风险时,临床医生更可能实施预防性治疗,从而改变了算法原本要预测的目标。一种简单的方法是忽略CMI,仅监测未接受治疗的患者(其结局未受影响)。一般而言,忽略CMI可能增加第一类错误,原因在于(i)未治疗患者不成比例地代表低预测风险人群,(ii)模型本身及临床医生对模型信任度的演变可能引发违反标准假设的复杂依赖关系。尽管如此,我们证明若监测条件性能,且条件可交换性或时不变选择偏倚成立,仍能进行有效推断。具体而言,我们开发了一种具有动态控制限的新型评分累积和(CUSUM)监测程序。通过模拟实验,我们展示了将模型更新与监测相结合的益处,并探究了对预测模型的过度信任如何延迟性能恶化的检测。最后,我们以COVID-19大流行期间术后恶心呕吐ML风险计算器的校准衰减检测为例,说明这些监测方法的应用。