Machine learning models are increasingly used in high-stakes decision-making systems. In such applications, a major concern is that these models sometimes discriminate against certain demographic groups such as individuals with certain race, gender, or age. Another major concern in these applications is the violation of the privacy of users. While fair learning algorithms have been developed to mitigate discrimination issues, these algorithms can still leak sensitive information, such as individuals' health or financial records. Utilizing the notion of differential privacy (DP), prior works aimed at developing learning algorithms that are both private and fair. However, existing algorithms for DP fair learning are either not guaranteed to converge or require full batch of data in each iteration of the algorithm to converge. In this paper, we provide the first stochastic differentially private algorithm for fair learning that is guaranteed to converge. Here, the term "stochastic" refers to the fact that our proposed algorithm converges even when minibatches of data are used at each iteration (i.e. stochastic optimization). Our framework is flexible enough to permit different fairness notions, including demographic parity and equalized odds. In addition, our algorithm can be applied to non-binary classification tasks with multiple (non-binary) sensitive attributes. As a byproduct of our convergence analysis, we provide the first utility guarantee for a DP algorithm for solving nonconvex-strongly concave min-max problems. Our numerical experiments show that the proposed algorithm consistently offers significant performance gains over the state-of-the-art baselines, and can be applied to larger scale problems with non-binary target/sensitive attributes.
翻译:机器学习模型越来越多地应用于高风险决策系统中。在此类应用中,一个主要担忧是这些模型有时会歧视某些人口群体,例如特定种族、性别或年龄的个体。另一个主要担忧是用户隐私的侵犯。尽管已开发出公平学习算法以缓解歧视问题,但这些算法仍可能泄露敏感信息,例如个人健康或财务记录。利用差分隐私的概念,先前工作旨在开发兼具隐私保护与公平性的学习算法。然而,现有的差分隐私公平学习算法要么无法保证收敛,要么需要在每次算法迭代中使用完整数据批次才能收敛。本文首次提出一种保证收敛的随机差分隐私公平学习算法。此处术语“随机”指代我们提出的算法即使每次迭代使用小批量数据(即随机优化)也能收敛。我们的框架足够灵活,可兼容不同公平性概念,包括人口统计均等和均等机会。此外,该算法还可应用于包含多个(非二元)敏感属性的非二元分类任务。作为收敛分析的副产品,我们首次为求解非凸-强凹极小极大问题的差分隐私算法提供了效用保证。数值实验表明,该算法相较于最先进的基线方法始终展现出显著的性能提升,并可应用于更大规模的非二元目标/敏感属性问题。