Predictive risk models in the public sector are commonly developed using administrative data that is more complete for subpopulations that more greatly rely on public services. In the United States, for instance, information on health care utilization is routinely available to government agencies for individuals supported by Medicaid and Medicare, but not for the privately insured. Critiques of public sector algorithms have identified such differential feature under-reporting as a driver of disparities in algorithmic decision-making. Yet this form of data bias remains understudied from a technical viewpoint. While prior work has examined the fairness impacts of additive feature noise and features that are clearly marked as missing, the setting of data missingness absent indicators (i.e. differential feature under-reporting) has been lacking in research attention. In this work, we present an analytically tractable model of differential feature under-reporting which we then use to characterize the impact of this kind of data bias on algorithmic fairness. We demonstrate how standard missing data methods typically fail to mitigate bias in this setting, and propose a new set of methods specifically tailored to differential feature under-reporting. Our results show that, in real world data settings, under-reporting typically leads to increasing disparities. The proposed solution methods show success in mitigating increases in unfairness.
翻译:公共部门的预测风险模型通常利用行政数据开发,这些数据对更依赖公共服务的子群体更为完整。例如在美国,政府机构通常可获得享受医疗补助和医疗保险的个人医疗利用信息,但无法获取私人保险患者的相关数据。对公共部门算法的批评指出,这种差异化特征少报是导致算法决策差异化的因素之一。然而从技术角度看,此类数据偏差仍未得到充分研究。尽管先前研究已探讨了加性特征噪声和明确标记缺失特征对公平性的影响,但缺乏指示符的数据缺失场景(即差异化特征少报)尚未受到足够关注。本研究提出一个可解析分析的差异化特征少报模型,并据此刻画此类数据偏差对算法公平性的影响。我们论证了标准缺失数据处理方法在此场景下通常无法缓解偏差,并提出一组专门针对差异化特征少报的新方法。结果表明,在真实数据场景中,少报通常会导致差异扩大,而所提出的解决方案能有效缓解不公平性的加剧。