This paper studies algorithmic fairness when the protected attribute is location. To handle protected attributes that are continuous, such as age or income, the standard approach is to discretize the domain into predefined groups, and compare algorithmic outcomes across groups. However, applying this idea to location raises concerns of gerrymandering and may introduce statistical bias. Prior work addresses these concerns but only for regularly spaced locations, while raising other issues, most notably its inability to discern regions that are likely to exhibit spatial unfairness. Similar to established notions of algorithmic fairness, we define spatial fairness as the statistical independence of outcomes from location. This translates into requiring that for each region of space, the distribution of outcomes is identical inside and outside the region. To allow for localized discrepancies in the distribution of outcomes, we compare how well two competing hypotheses explain the observed outcomes. The null hypothesis assumes spatial fairness, while the alternate allows different distributions inside and outside regions. Their goodness of fit is then assessed by a likelihood ratio test. If there is no significant difference in how well the two hypotheses explain the observed outcomes, we conclude that the algorithm is spatially fair.
翻译:本文研究当受保护属性为地理位置时的算法公平性问题。针对年龄或收入等连续型受保护属性,标准方法是将域离散化为预定义组别,并比较各组间的算法结果。然而,将该方法应用于地理位置会引发选区划分操纵(gerrymandering)问题,并可能引入统计偏差。现有研究虽能解决上述问题,但仅适用于规则间隔的地理位置,同时产生其他问题,最显著的是无法识别可能呈现空间不公平性的区域。与算法公平性的既定概念类似,本文将空间公平性定义为结果变量与地理位置的统计独立性。这转化为要求:对于每个空间区域,结果变量在区域内部与外部的分布完全相同。为允许结果分布存在局部差异,我们比较两个竞争性假设对观测结果的解释能力。零假设假定空间公平性成立,而备择假设允许区域内外存在不同分布。随后通过似然比检验评估两者的拟合优度。若两个假设解释观测结果的能力无显著差异,则判定该算法具有空间公平性。