Electronic health records (EHRs) serve as an essential data source for the envisioned artificial intelligence (AI)-driven transformation in healthcare. However, clinician biases reflected in EHR notes can lead to AI models inheriting and amplifying these biases, perpetuating health disparities. This study investigates the impact of stigmatizing language (SL) in EHR notes on mortality prediction using a Transformer-based deep learning model and explainable AI (XAI) techniques. Our findings demonstrate that SL written by clinicians adversely affects AI performance, particularly so for black patients, highlighting SL as a source of racial disparity in AI model development. To explore an operationally efficient way to mitigate SL's impact, we investigate patterns in the generation of SL through a clinicians' collaborative network, identifying central clinicians as having a stronger impact on racial disparity in the AI model. We find that removing SL written by central clinicians is a more efficient bias reduction strategy than eliminating all SL in the entire corpus of data. This study provides actionable insights for responsible AI development and contributes to understanding clinician behavior and EHR note writing in healthcare.
翻译:电子健康记录(EHR)是推动医疗领域人工智能(AI)变革所设想的关键数据来源。然而,EHR笔记中反映的临床医生偏见可能导致AI模型继承并放大这些偏见,从而加剧健康差异。本研究基于Transformer深度学习模型和可解释AI(XAI)技术,探讨了EHR笔记中污名化语言(SL)对死亡率预测的影响。我们的发现表明,临床医生书写的SL会负面影响AI性能,尤其是在黑人患者中更为显著,凸显了SL作为AI模型开发中种族差异的一个来源。为探索缓解SL影响的操作高效方法,我们通过临床医生协作网络研究了SL的产生模式,识别出核心临床医生对AI模型中的种族差异具有更强的影响力。研究发现,仅移除核心临床医生所书写的SL比消除整个数据集中所有SL更能实现高效的偏差减少策略。本研究为负责任的AI开发提供了可操作见解,并有助于理解临床医生行为及医疗领域的EHR笔记书写。