Widespread adoption of AI for medical decision making is still hindered due to ethical and safety-related concerns. For AI-based decision support systems in healthcare settings it is paramount to be reliable and trustworthy. Common deep learning approaches, however, have the tendency towards overconfidence under data shift. Such inappropriate extrapolation beyond evidence-based scenarios may have dire consequences. This highlights the importance of reliable estimation of local uncertainty and its communication to the end user. While stochastic neural networks have been heralded as a potential solution to these issues, this study investigates their actual reliability in clinical applications. We centered our analysis on the exemplary use case of mortality prediction for ICU hospitalizations using EHR from MIMIC3 study. For predictions on the EHR time series, Encoder-Only Transformer models were employed. Stochasticity of model functions was achieved by incorporating common methods such as Bayesian neural network layers and model ensembles. Our models achieve state of the art performance in terms of discrimination performance (AUC ROC: 0.868+-0.011, AUC PR: 0.554+-0.034) and calibration on the mortality prediction benchmark. However, epistemic uncertainty is critically underestimated by the selected stochastic deep learning methods. A heuristic proof for the responsible collapse of the posterior distribution is provided. Our findings reveal the inadequacy of commonly used stochastic deep learning approaches to reliably recognize OoD samples. In both methods, unsubstantiated model confidence is not prevented due to strongly biased functional posteriors, rendering them inappropriate for reliable clinical decision support. This highlights the need for approaches with more strictly enforced or inherent distance-awareness to known data points, e.g., using kernel-based techniques.
翻译:人工智能在医疗决策中的广泛应用仍受制于伦理和安全性考量。在医疗场景中,基于AI的决策支持系统必须兼具可靠性和可信赖性。然而,常见深度学习方法在数据偏移下存在过度自信的倾向。这种超出循证场景的不当外推可能带来严重后果。这凸显了可靠估计局部不确定性并将其传达给最终用户的重要性。尽管随机神经网络被视作解决这些问题的潜在方案,本研究探究了其在临床应用中的实际可靠性。我们以MIMIC3研究中的电子健康记录(EHR)预测ICU住院患者死亡率为典型案例展开分析。针对EHR时间序列的预测任务,采用编码器-仅Transformer模型。通过纳入贝叶斯神经网络层和模型集成等常见方法实现模型函数的随机性。我们的模型在死亡率预测基准上取得了判别性能(AUC ROC:0.868±0.011,AUC PR:0.554±0.034)与校准能力的最优表现。然而,所选随机深度学习方法显著低估了认知不确定性。我们提供了后验分布崩溃的启发式证明。研究结果表明,常用随机深度学习方法无法可靠识别分布外(OoD)样本。两种方法中,因功能后验存在严重偏差,均未能防止无根据的模型过度自信,使其不适用于可靠临床决策支持。这凸显了采用更严格约束或固有距离感知机制(如基于核的技术)对已知数据点进行处理的必要性。