Demographic parity is the most widely recognized measure of group fairness in machine learning, which ensures equal treatment of different demographic groups. Numerous works aim to achieve demographic parity by pursuing the commonly used metric $\Delta DP$. Unfortunately, in this paper, we reveal that the fairness metric $\Delta DP$ can not precisely measure the violation of demographic parity, because it inherently has the following drawbacks: \textit{i)} zero-value $\Delta DP$ does not guarantee zero violation of demographic parity, \textit{ii)} $\Delta DP$ values can vary with different classification thresholds. To this end, we propose two new fairness metrics, \textsf{A}rea \textsf{B}etween \textsf{P}robability density function \textsf{C}urves (\textsf{ABPC}) and \textsf{A}rea \textsf{B}etween \textsf{C}umulative density function \textsf{C}urves (\textsf{ABCC}), to precisely measure the violation of demographic parity in distribution level. The new fairness metrics directly measure the difference between the distributions of the prediction probability for different demographic groups. Thus our proposed new metrics enjoy: \textit{i)} zero-value \textsf{ABCC}/\textsf{ABPC} guarantees zero violation of demographic parity; \textit{ii)} \textsf{ABCC}/\textsf{ABPC} guarantees demographic parity while the classification threshold adjusted. We further re-evaluate the existing fair models with our proposed fairness metrics and observe different fairness behaviors of those models under the new metrics.
翻译:人口均等是机器学习中最广泛认可的群体公平性度量标准,旨在确保不同人口群体受到平等对待。众多研究致力于通过追求常用的$\Delta DP$指标来实现人口均等。然而,本文揭示公平性指标$\Delta DP$无法精确衡量人口均等的违背程度,因其固有缺陷:\textit{i)} 零值的$\Delta DP$无法保证人口均等零违背;\textit{ii)} $\Delta DP$值会随分类阈值变化而变化。为此,我们提出两种新的公平性指标——概率密度函数曲线间面积(\textsf{ABPC})和累积密度函数曲线间面积(\textsf{ABCC}),以在分布层面精确度量人口均等的违背程度。新指标直接度量不同人口群体预测概率分布之间的差异。因此,本文提出的新指标具有以下优势:\textit{i)} 零值的\textsf{ABCC}/\textsf{ABPC}保证人口均等零违背;\textit{ii)} 调整分类阈值时,\textsf{ABCC}/\textsf{ABPC}仍能保证人口均等。我们进一步基于新指标重新评估现有公平模型,观察到这些模型在新指标下的公平性行为存在差异。