The vast majority of techniques to train fair models require access to the protected attribute (e.g., race, gender), either at train time or in production. However, in many important applications this protected attribute is largely unavailable. In this paper, we develop methods for measuring and reducing fairness violations in a setting with limited access to protected attribute labels. Specifically, we assume access to protected attribute labels on a small subset of the dataset of interest, but only probabilistic estimates of protected attribute labels (e.g., via Bayesian Improved Surname Geocoding) for the rest of the dataset. With this setting in mind, we propose a method to estimate bounds on common fairness metrics for an existing model, as well as a method for training a model to limit fairness violations by solving a constrained non-convex optimization problem. Unlike similar existing approaches, our methods take advantage of contextual information -- specifically, the relationships between a model's predictions and the probabilistic prediction of protected attributes, given the true protected attribute, and vice versa -- to provide tighter bounds on the true disparity. We provide an empirical illustration of our methods using voting data. First, we show our measurement method can bound the true disparity up to 5.5x tighter than previous methods in these applications. Then, we demonstrate that our training technique effectively reduces disparity while incurring lesser fairness-accuracy trade-offs than other fair optimization methods with limited access to protected attributes.
翻译:绝大多数训练公平模型的技术需要在训练时或生产环境中访问受保护属性(例如人种、性别)。然而,在许多重要应用中,这种受保护属性通常难以获取。本文针对受保护属性标签访问受限的场景,开发了用于测量和减少公平性违规的方法。具体而言,我们假设仅能获取数据集小部分样本的受保护属性标签,其余样本只能获得受保护属性标签的概率估计(例如通过贝叶斯改进姓氏地理编码)。基于这一设定,我们提出了一种估计现有模型常见公平性指标边界的方法,以及通过求解约束非凸优化问题来训练模型以限制公平性违规的方法。与现有类似方法不同,我们的方法利用上下文信息——特别是模型预测与受保护属性概率预测之间在真实受保护属性条件下的关系,反之亦然——从而为真实差异提供更紧凑的边界。我们使用投票数据进行了实证说明:首先证明我们的测量方法在这些应用场景中能将真实差异边界压缩至以往方法的5.5倍以内;其次验证了我们的训练技术能在受保护属性访问受限的情况下有效降低差异,同时比同类公平优化方法产生更小的公平性-准确性权衡损失。