Clinical artificial intelligence (AI) systems routinely produce predictions without principled quantification of uncertainty, limiting their trustworthiness in high-stakes medical environments. This paper presents an integrated research programme addressing two interconnected problems: (1) the development of a fully end-to-end Bayesian uncertainty modelling framework for multimodal clinical data, and (2) the application of calibrated uncertainty estimates as a formal measure of algorithmic equity across patient subgroups. We construct a probabilistic deep learning architecture comprising modality-specific variational encoders, a precision-weighted late fusion mechanism, and a decomposed uncertainty output head that separates aleatoric from epistemic uncertainty. The system is trained with a composite Bayesian loss incorporating binary cross-entropy, Kullback-Leibler divergence regularisation, and an uncertainty calibration penalty. We evaluate model calibration using Expected Calibration Error (ECE = 0.096) and conduct a subgroup equity audit across facility type, socioeconomic status, age group, and biological sex on a dataset of 1,000 simulated patients. Results demonstrate that epistemic uncertainty systematically identifies underserved populations: primary/rural facility patients show a 15.3% uncertainty equity gap (p < 0.001, effect size = 0.698), low socioeconomic status patients exhibit a 6.8% gap (p < 0.001), and elderly patients show a 3.9% gap (p < 0.001), whilst no significant sex-based disparity is detected. These findings establish that calibrated uncertainty is not merely a technical property of probabilistic models but constitutes an actionable equity signal with direct clinical relevance.
翻译:临床人工智能系统在预测过程中普遍缺乏对不确定性的原则性量化,这限制了其在高风险医疗环境中的可信度。本文提出一项整合研究方案,旨在解决两个相互关联的问题:(1) 针对多模态临床数据开发完全端到端的贝叶斯不确定性建模框架,以及(2) 将校准后的不确定性估计作为患者亚组间算法公平性的形式化度量。我们构建了一个概率深度学习架构,包含模态特异性变分编码器、精度加权晚期融合机制,以及分离偶然不确定性与认知不确定性的分解不确定性输出头。该系统采用复合贝叶斯损失进行训练,该损失函数整合了二元交叉熵、Kullback-Leibler散度正则化及不确定性校准惩罚项。我们使用期望校准误差(ECE=0.096)评估模型校准性能,并在包含1000名模拟患者的数据集上,针对设施类型、社会经济地位、年龄组和生物学性别进行亚组公平性审计。结果表明,认知不确定性能够系统性识别服务不足人群:基层/农村设施患者的不确定性公平性差距达15.3%(p<0.001,效应量=0.698),低社会经济地位患者差距为6.8%(p<0.001),老年患者为3.9%(p<0.001),同时未检测到显著的性别差异。这些发现证实,校准后的不确定性不仅是概率模型的技术属性,更是具有直接临床相关性的可操作公平性信号。