In machine learning fairness, training models that minimize disparity across different sensitive groups often leads to diminished accuracy, a phenomenon known as the fairness-accuracy trade-off. The severity of this trade-off inherently depends on dataset characteristics such as dataset imbalances or biases and therefore, using a uniform fairness requirement across diverse datasets remains questionable. To address this, we present a computationally efficient approach to approximate the fairness-accuracy trade-off curve tailored to individual datasets, backed by rigorous statistical guarantees. By utilizing the You-Only-Train-Once (YOTO) framework, our approach mitigates the computational burden of having to train multiple models when approximating the trade-off curve. Crucially, we introduce a novel methodology for quantifying uncertainty in our estimates, thereby providing practitioners with a robust framework for auditing model fairness while avoiding false conclusions due to estimation errors. Our experiments spanning tabular (e.g., Adult), image (CelebA), and language (Jigsaw) datasets underscore that our approach not only reliably quantifies the optimum achievable trade-offs across various data modalities but also helps detect suboptimality in SOTA fairness methods.
翻译:在机器学习公平性研究中,训练能够最小化不同敏感群体间差异的模型通常会导致准确率下降,这一现象被称为公平性与准确率的权衡。这种权衡的严重程度本质上取决于数据集的特征,例如数据集的不平衡性或偏差,因此,在不同数据集上采用统一的公平性要求仍然值得商榷。为解决此问题,我们提出了一种计算高效的方法,用以近似拟合针对特定数据集的公平性-准确率权衡曲线,该方法具有严格的统计保证。通过利用"仅需训练一次"(YOTO)框架,我们的方法减轻了在近似权衡曲线时需要训练多个模型所带来的计算负担。至关重要的是,我们引入了一种量化估计不确定性的新方法,从而为实践者提供了一个鲁棒的框架,用于审计模型公平性,同时避免因估计误差而得出错误结论。我们在表格数据(如Adult)、图像数据(CelebA)和语言数据(Jigsaw)数据集上进行的实验表明,我们的方法不仅能够可靠地量化跨多种数据模态可实现的最佳权衡,还有助于检测当前最先进(SOTA)公平性方法中的次优性。