A general belief in fair classification is that fairness constraints incur a trade-off with accuracy, which biased data may worsen. Contrary to this belief, Blum & Stangl (2019) show that fair classification with equal opportunity constraints even on extremely biased data can recover optimally accurate and fair classifiers on the original data distribution. Their result is interesting because it demonstrates that fairness constraints can implicitly rectify data bias and simultaneously overcome a perceived fairness-accuracy trade-off. Their data bias model simulates under-representation and label bias in underprivileged population, and they show the above result on a stylized data distribution with i.i.d. label noise, under simple conditions on the data distribution and bias parameters. We propose a general approach to extend the result of Blum & Stangl (2019) to different fairness constraints, data bias models, data distributions, and hypothesis classes. We strengthen their result, and extend it to the case when their stylized distribution has labels with Massart noise instead of i.i.d. noise. We prove a similar recovery result for arbitrary data distributions using fair reject option classifiers. We further generalize it to arbitrary data distributions and arbitrary hypothesis classes, i.e., we prove that for any data distribution, if the optimally accurate classifier in a given hypothesis class is fair and robust, then it can be recovered through fair classification with equal opportunity constraints on the biased distribution whenever the bias parameters satisfy certain simple conditions. Finally, we show applications of our technique to time-varying data bias in classification and fair machine learning pipelines.
翻译:一个普遍的认知是,公平分类会引入准确率与公平性之间的权衡,而有偏数据可能加剧这一权衡。与这一认知相悖的是,Blum和Stangl(2019)的研究表明,即使在极端有偏数据上施加机会均等约束的公平分类,也能在原始数据分布上恢复最优的准确率与公平性分类器。这一结果引人关注,因为它揭示了公平性约束能够隐式纠正数据偏差,同时克服看似存在的公平-准确率权衡。他们的数据偏差模型模拟了弱势群体的代表性不足和标签偏差,并在满足数据分布与偏差参数简单条件的、带有独立同分布标签噪声的典型数据分布下证明了上述结果。我们提出了一个通用方法,将Blum和Stangl(2019)的结论推广至不同的公平性约束、数据偏差模型、数据分布和假设类别。我们强化了他们的结论,并将其扩展至其典型分布中的标签带有Massart噪声(而非独立同分布噪声)的情形。我们利用公平拒绝选项分类器,对任意数据分布证明了类似的恢复结论。随后,我们将其进一步推广至任意数据分布和任意假设类别,即证明:对于任意数据分布,若给定假设类别中的最优准确率分类器是公平且鲁棒的,则当偏差参数满足某些简单条件时,可以通过对有偏分布施加机会均等约束的公平分类来恢复该分类器。最后,我们展示了该方法在分类和公平机器学习流水线中时变数据偏差问题上的应用。