Electronic health records (EHRs) serve as an essential data source for the envisioned artificial intelligence (AI)-driven transformation in healthcare. However, clinician biases reflected in EHR notes can lead to AI models inheriting and amplifying these biases, perpetuating health disparities. This study investigates the impact of stigmatizing language (SL) in EHR notes on mortality prediction using a Transformer-based deep learning model and explainable AI (XAI) techniques. Our findings demonstrate that SL written by clinicians adversely affects AI performance, particularly so for black patients, highlighting SL as a source of racial disparity in AI model development. To explore an operationally efficient way to mitigate SL's impact, we investigate patterns in the generation of SL through a clinicians' collaborative network, identifying central clinicians as having a stronger impact on racial disparity in the AI model. We find that removing SL written by central clinicians is a more efficient bias reduction strategy than eliminating all SL in the entire corpus of data. This study provides actionable insights for responsible AI development and contributes to understanding clinician behavior and EHR note writing in healthcare.
翻译:电子健康记录(EHRs)是设想中人工智能(AI)驱动医疗变革的重要数据来源。然而,EHR记录中反映的临床医生偏见可能导致AI模型继承并放大这些偏见,从而加剧健康差异。本研究通过基于Transformer的深度学习模型与可解释人工智能(XAI)技术,探究EHR记录中的污名化语言(SL)对死亡率预测的影响。研究发现,临床医生使用的SL会显著降低AI性能(尤其对黑人患者而言),揭示SL是AI模型开发中种族差异的来源之一。为探索减轻SL影响的高效操作方案,我们通过临床医生协作网络分析了SL的生成模式,发现核心临床医生对AI模型中的种族差异影响更大。结果表明,移除核心临床医生书写的SL比消除数据全集中所有SL能更有效地减少偏见。本研究为负责任AI开发提供可操作见解,并有助于理解医疗场景中临床医生行为与EHR记录撰写。