Artificial intelligence (AI) models trained using medical images for clinical tasks often exhibit bias in the form of disparities in performance between subgroups. Since not all sources of biases in real-world medical imaging data are easily identifiable, it is challenging to comprehensively assess how those biases are encoded in models, and how capable bias mitigation methods are at ameliorating performance disparities. In this article, we introduce a novel analysis framework for systematically and objectively investigating the impact of biases in medical images on AI models. We developed and tested this framework for conducting controlled in silico trials to assess bias in medical imaging AI using a tool for generating synthetic magnetic resonance images with known disease effects and sources of bias. The feasibility is showcased by using three counterfactual bias scenarios to measure the impact of simulated bias effects on a convolutional neural network (CNN) classifier and the efficacy of three bias mitigation strategies. The analysis revealed that the simulated biases resulted in expected subgroup performance disparities when the CNN was trained on the synthetic datasets. Moreover, reweighing was identified as the most successful bias mitigation strategy for this setup, and we demonstrated how explainable AI methods can aid in investigating the manifestation of bias in the model using this framework. Developing fair AI models is a considerable challenge given that many and often unknown sources of biases can be present in medical imaging datasets. In this work, we present a novel methodology to objectively study the impact of biases and mitigation strategies on deep learning pipelines, which can support the development of clinical AI that is robust and responsible.
翻译:人工智能模型在医学影像驱动的临床任务中,常因亚组间性能差异而表现出偏差。由于真实世界医学影像数据中的偏差来源难以全面识别,如何系统评估模型编码偏差的程度及现有偏差缓解方法对性能差异的改善效果成为重大挑战。本文提出一种新型分析框架,用于系统且客观地研究医学影像偏差对AI模型的影响。我们开发并测试了该框架,通过生成已知疾病效应和偏差源的合成磁共振图像工具,开展受控计算机模拟试验以实现医学影像AI偏差评估。通过三种反事实偏差场景验证可行性,分别评估模拟偏差对卷积神经网络分类器的影响及三种偏差缓解策略的有效性。分析表明,当使用合成数据集训练CNN时,模拟偏差会导致预期的亚组性能差异。此外,重加权被确定为该设置下最成功的偏差缓解策略,同时我们展示了可解释AI方法如何借助该框架辅助研究模型偏差表现。鉴于医学影像数据集中普遍存在大量未知偏差源,开发公平AI模型仍是重大挑战。本研究提出了一种新方法学,可客观研究偏差及缓解策略对深度学习流水线的影响,为开发稳健且负责任的临床AI提供支持。