Widespread adoption of AI for medical decision making is still hindered due to ethical and safety-related concerns. For AI-based decision support systems in healthcare settings it is paramount to be reliable and trustworthy. Common deep learning approaches, however, have the tendency towards overconfidence under data shift. Such inappropriate extrapolation beyond evidence-based scenarios may have dire consequences. This highlights the importance of reliable estimation of local uncertainty and its communication to the end user. While stochastic neural networks have been heralded as a potential solution to these issues, this study investigates their actual reliability in clinical applications. We centered our analysis on the exemplary use case of mortality prediction for ICU hospitalizations using EHR from MIMIC3 study. For predictions on the EHR time series, Encoder-Only Transformer models were employed. Stochasticity of model functions was achieved by incorporating common methods such as Bayesian neural network layers and model ensembles. Our models achieve state of the art performance in terms of discrimination performance (AUC ROC: 0.868+-0.011, AUC PR: 0.554+-0.034) and calibration on the mortality prediction benchmark. However, epistemic uncertainty is critically underestimated by the selected stochastic deep learning methods. A heuristic proof for the responsible collapse of the posterior distribution is provided. Our findings reveal the inadequacy of commonly used stochastic deep learning approaches to reliably recognize OoD samples. In both methods, unsubstantiated model confidence is not prevented due to strongly biased functional posteriors, rendering them inappropriate for reliable clinical decision support. This highlights the need for approaches with more strictly enforced or inherent distance-awareness to known data points, e.g., using kernel-based techniques.
翻译:人工智能在医疗决策中的广泛应用仍受伦理和安全性问题的阻碍。对于医疗环境中的基于AI的决策支持系统,可靠性和可信赖性至关重要。然而,常见的深度学习方法在数据偏移下倾向于过度自信。这种超出基于证据场景的不当外推可能导致严重后果,凸显了可靠估计局部不确定性并将其传达给最终用户的重要性。尽管随机神经网络被认为是解决这些问题的潜在方案,本研究调查了它们在临床应用中的实际可靠性。我们以MIMIC3研究中电子健康记录(EHR)数据的ICU住院患者死亡率预测作为典型案例进行分析。针对EHR时间序列预测,我们采用了编码器-仅Transformer模型。通过引入贝叶斯神经网络层和模型集成等常见方法实现了模型函数的随机性。我们的模型在死亡率预测基准上达到了最先进的判别性能(AUC ROC: 0.868±0.011, AUC PR: 0.554±0.034)和校准能力。然而,所选随机深度学习方法严重低估了认知不确定性。本文提供了后验分布崩溃的启发式证明。研究发现,常用随机深度学习方法在可靠识别分布外样本方面存在不足。两种方法中,由于高度偏差的函数后验分布,未经验证的模型置信度无法避免,使其不适用于可靠的临床决策支持。这凸显了采用更严格约束或固有距离感知方法(如基于核的技术)处理已知数据点的必要性。