In machine learning fairness, training models which minimize disparity across different sensitive groups often leads to diminished accuracy, a phenomenon known as the fairness-accuracy trade-off. The severity of this trade-off fundamentally depends on dataset characteristics such as dataset imbalances or biases. Therefore using a uniform fairness requirement across datasets remains questionable and can often lead to models with substantially low utility. To address this, we present a computationally efficient approach to approximate the fairness-accuracy trade-off curve tailored to individual datasets, backed by rigorous statistical guarantees. By utilizing the You-Only-Train-Once (YOTO) framework, our approach mitigates the computational burden of having to train multiple models when approximating the trade-off curve. Moreover, we quantify the uncertainty in our approximation by introducing confidence intervals around this curve, offering a statistically grounded perspective on the acceptable range of fairness violations for any given accuracy threshold. Our empirical evaluation spanning tabular, image and language datasets underscores that our approach provides practitioners with a principled framework for dataset-specific fairness decisions across various data modalities.
翻译:在机器学习公平性中,训练模型以最小化不同敏感群体间的差异往往会导致准确率下降,即公平性-准确率权衡现象。该权衡的严重程度根本上取决于数据集特征(如数据不平衡或偏差)。因此,对数据集采用统一的公平性要求仍存疑义,且常导致模型效用显著降低。为解决此问题,我们提出了一种计算高效的方法,可针对个体数据集近似公平性-准确率权衡曲线,并辅以严格的统计保证。通过利用YOTO(仅需单次训练)框架,我们的方法可缓解近似权衡曲线时需训练多个模型的计算负担。此外,我们通过引入该曲线的置信区间来量化近似的不确定性,为任意给定准确率阈值下可接受的公平性违规范围提供了统计视角。涵盖表格、图像和语言数据集的实证评估表明,我们的方法为从业者提供了跨数据模态的、基于数据集特征的公平性决策原则性框架。