Algorithmic decision making driven by neural networks has become very prominent in applications that directly affect people's quality of life. In this paper, we study the problem of verifying, training, and guaranteeing individual fairness of neural network models. A popular approach for enforcing fairness is to translate a fairness notion into constraints over the parameters of the model. However, such a translation does not always guarantee fair predictions of the trained neural network model. To address this challenge, we develop a counterexample-guided post-processing technique to provably enforce fairness constraints at prediction time. Contrary to prior work that enforces fairness only on points around test or train data, we are able to enforce and guarantee fairness on all points in the input domain. Additionally, we propose an in-processing technique to use fairness as an inductive bias by iteratively incorporating fairness counterexamples in the learning process. We have implemented these techniques in a tool called FETA. Empirical evaluation on real-world datasets indicates that FETA is not only able to guarantee fairness on-the-fly at prediction time but also is able to train accurate models exhibiting a much higher degree of individual fairness.
翻译:基于神经网络的算法决策已深刻融入直接影响人们生活质量的应用领域。本文研究了神经网络模型的个体公平性验证、训练与保障问题。当前一种主流的公平性实施方法是将公平性概念转化为模型参数上的约束条件,然而此类转化方法无法保证训练后的神经网络模型始终输出公平的预测结果。为应对这一挑战,我们提出了一种基于反例引导的后处理技术,能够在预测阶段可证明地强制满足公平性约束。与现有仅在测试或训练数据邻域内实施公平性的工作不同,我们的方法能够在输入域内所有数据点上强制并保证公平性。此外,我们还提出一种内处理技术,通过迭代地将公平性反例融入学习过程,将公平性作为归纳偏置进行训练。这些技术已在一个名为FETA的工具中实现。在真实数据集上的实验评估表明,FETA不仅能在预测阶段实时保证公平性,还能训练出具有更高个体公平性程度的精准模型。