Class imbalance and group (e.g., race, gender, and age) imbalance are acknowledged as two reasons in data that hinder the trade-off between fairness and utility of machine learning classifiers. Existing techniques have jointly addressed issues regarding class imbalance and group imbalance by proposing fair over-sampling techniques. Unlike the common oversampling techniques, which only address class imbalance, fair oversampling techniques significantly improve the abovementioned trade-off, as they can also address group imbalance. However, if the size of the original clusters is too small, these techniques may cause classifier overfitting. To address this problem, we herein develop a fair oversampling technique using data from heterogeneous clusters. The proposed technique generates synthetic data that have class-mix features or group-mix features to make classifiers robust to overfitting. Moreover, we develop an interpolation method that can enhance the validity of generated synthetic data by considering the original cluster distribution and data noise. Finally, we conduct experiments on five realistic datasets and three classifiers, and the experimental results demonstrate the effectiveness of the proposed technique in terms of fairness and utility.
翻译:类别不平衡和群体(如种族、性别和年龄)不平衡被认为是数据中阻碍机器学习分类器公平性与效用性权衡的两个原因。现有技术通过提出公平过采样技术,共同解决了类别不平衡和群体不平衡的问题。与仅处理类别不平衡的常见过采样技术不同,公平过采样技术显著改善了上述权衡,因为它们还能处理群体不平衡。然而,如果原始聚类规模过小,这些技术可能导致分类器过拟合。为解决此问题,本文开发了一种使用异质聚类数据的公平过采样技术。所提出的技术生成具有类别混合特征或群体混合特征的合成数据,使分类器对过拟合具有鲁棒性。此外,我们开发了一种插值方法,通过考虑原始聚类分布和数据噪声,增强生成合成数据的有效性。最后,我们在五个真实数据集和三个分类器上进行了实验,实验结果证明了所提出技术在公平性和效用性方面的有效性。