Advances in artificial intelligence (AI) have achieved expert-level performance in medical imaging applications. Notably, self-supervised vision-language foundation models can detect a broad spectrum of pathologies without relying on explicit training annotations. However, it is crucial to ensure that these AI models do not mirror or amplify human biases, thereby disadvantaging historically marginalized groups such as females or Black patients. The manifestation of such biases could systematically delay essential medical care for certain patient subgroups. In this study, we investigate the algorithmic fairness of state-of-the-art vision-language foundation models in chest X-ray diagnosis across five globally-sourced datasets. Our findings reveal that compared to board-certified radiologists, these foundation models consistently underdiagnose marginalized groups, with even higher rates seen in intersectional subgroups, such as Black female patients. Such demographic biases present over a wide range of pathologies and demographic attributes. Further analysis of the model embedding uncovers its significant encoding of demographic information. Deploying AI systems with these biases in medical imaging can intensify pre-existing care disparities, posing potential challenges to equitable healthcare access and raising ethical questions about their clinical application.
翻译:人工智能(AI)的进步已使医学影像应用达到专家级性能。值得注意的是,自监督视觉语言基础模型无需依赖显式训练标注就能检测多种病理特征。然而,确保这些AI模型不反映或放大人类偏见,从而避免对历史上被边缘化的群体(如女性患者或黑人患者)造成不利影响至关重要。此类偏见的显现可能系统性地延误特定患者群体的必要医疗护理。本研究通过五个全球来源的数据集,考察了胸部X射线诊断中最先进的视觉语言基础模型的算法公平性。研究结果表明,与获得专业认证的放射科医师相比,这些基础模型始终对边缘化群体存在诊断不足的问题,而在交叉群体(如黑人女性患者)中诊断不足率更高。这类人口统计学偏见广泛存在于多种病理特征和人口属性中。对模型嵌入的进一步分析发现,其显著编码了人口统计学信息。在医学影像中部署带有这些偏见的AI系统可能加剧既有的医疗差距,对公平获得医疗资源构成潜在挑战,并引发关于其临床应用的伦理争议。