Machine learning (ML) holds great promise for improving healthcare, but it is critical to ensure that its use will not propagate or amplify health disparities. An important step is to characterize the (un)fairness of ML models - their tendency to perform differently across subgroups of the population - and to understand its underlying mechanisms. One potential driver of algorithmic unfairness, shortcut learning, arises when ML models base predictions on improper correlations in the training data. However, diagnosing this phenomenon is difficult, especially when sensitive attributes are causally linked with disease. Using multi-task learning, we propose the first method to assess and mitigate shortcut learning as a part of the fairness assessment of clinical ML systems, and demonstrate its application to clinical tasks in radiology and dermatology. Finally, our approach reveals instances when shortcutting is not responsible for unfairness, highlighting the need for a holistic approach to fairness mitigation in medical AI.
翻译:机器学习(ML)在改善医疗保健方面具有巨大潜力,但确保其使用不会传播或加剧健康差异至关重要。一个关键步骤是描述机器学习模型的(不)公平性——即模型在不同人群子组中表现差异的倾向——并理解其潜在机制。算法不公平的一个潜在驱动因素是捷径学习,即当机器学习模型基于训练数据中的不当相关性进行预测时出现的问题。然而,诊断这一现象非常困难,尤其是在敏感属性与疾病存在因果关系的情况下。通过多任务学习,我们提出了第一种将捷径学习的评估和缓解作为临床机器学习系统公平性评估一部分的方法,并展示了其在放射学和皮肤病学临床任务中的应用。最后,我们的方法揭示了捷径学习并非不公平性原因的情况,突显了在医学AI中采取整体方法缓解不公平性的必要性。