Health care decisions are increasingly informed by clinical decision support algorithms, but these algorithms may perpetuate or increase racial and ethnic disparities in access to and quality of health care. Further complicating the problem, clinical data often have missing or poor quality racial and ethnic information, which can lead to misleading assessments of algorithmic bias. We present novel statistical methods that allow for the use of probabilities of racial/ethnic group membership in assessments of algorithm performance and quantify the statistical bias that results from error in these imputed group probabilities. We propose a sensitivity analysis approach to estimating the statistical bias that allows practitioners to assess disparities in algorithm performance under a range of assumed levels of group probability error. We also prove theoretical bounds on the statistical bias for a set of commonly used fairness metrics and describe real-world scenarios where our theoretical results are likely to apply. We present a case study using imputed race and ethnicity from the Bayesian Improved Surname Geocoding (BISG) algorithm for estimation of disparities in a clinical decision support algorithm used to inform osteoporosis treatment. Our novel methods allow policy makers to understand the range of potential disparities under a given algorithm even when race and ethnicity information is missing and to make informed decisions regarding the implementation of machine learning for clinical decision support.
翻译:医疗决策日益依赖临床决策支持算法,但这些算法可能延续或加剧医疗资源获取与质量中的种族及民族差异。更棘手的是,临床数据常存在种族和民族信息缺失或质量低下的问题,这可能导致对算法偏差的误导性评估。我们提出新的统计方法,允许在算法性能评估中使用种族/民族群体成员概率,并量化因这些估算群体概率误差导致的统计偏差。我们提出一种敏感性分析方法来估计统计偏差,使从业者能在不同假设的群体概率误差水平下评估算法性能差异。我们还针对一组常用公平性指标证明了统计偏差的理论界,并描述了理论结果可能适用的现实场景。通过贝叶斯改进姓氏地理编码(BISG)算法估算的种族和民族数据,我们展示了一项案例研究,评估了用于指导骨质疏松治疗的临床决策支持算法中的差异。我们的新方法使政策制定者能够在种族和民族信息缺失时,理解给定算法下潜在差异的范围,并就临床决策支持的机器学习实施做出明智决策。