Data practices shape research and practice on fairness in machine learning (fair ML). Critical data studies offer important reflections and critiques for the responsible advancement of the field by highlighting shortcomings and proposing recommendations for improvement. In this work, we present a comprehensive analysis of fair ML datasets, demonstrating how unreflective yet common practices hinder the reach and reliability of algorithmic fairness findings. We systematically study protected information encoded in tabular datasets and their usage in 280 experiments across 142 publications. Our analyses identify three main areas of concern: (1) a \textbf{lack of representation for certain protected attributes} in both data and evaluations; (2) the widespread \textbf{exclusion of minorities} during data preprocessing; and (3) \textbf{opaque data processing} threatening the generalization of fairness research. By conducting exemplary analyses on the utilization of prominent datasets, we demonstrate how unreflective data decisions disproportionately affect minority groups, fairness metrics, and resultant model comparisons. Additionally, we identify supplementary factors such as limitations in publicly available data, privacy considerations, and a general lack of awareness, which exacerbate these challenges. To address these issues, we propose a set of recommendations for data usage in fairness research centered on transparency and responsible inclusion. This study underscores the need for a critical reevaluation of data practices in fair ML and offers directions to improve both the sourcing and usage of datasets.
翻译:数据实践塑造了机器学习公平性(公平ML)的研究与实践。关键数据研究通过揭示不足并提出改进建议,为该领域的负责任发展提供了重要反思与批评。本文对公平ML数据集进行了全面分析,展示了未经反思的常见实践如何阻碍算法公平性发现的范围和可靠性。我们系统地研究了表格数据集中编码的保护性信息,及其在142篇论文的280项实验中的使用情况。分析发现三个主要问题领域:(1)数据和评估中**某些保护性属性代表性不足**;(2)数据预处理过程中普遍**排除少数群体**;(3)**不透明的数据处理**威胁公平性研究的泛化能力。通过对主要数据集使用情况的示例性分析,我们展示了未经反思的数据决策如何不成比例地影响少数群体、公平性指标及由此产生的模型比较。此外,我们识别出加剧这些挑战的辅助因素,包括公开数据获取限制、隐私考量及普遍的意识缺乏。为解决这些问题,我们提出了一套围绕透明性和负责任包容性的公平性研究数据使用建议。本研究强调了对公平ML数据实践进行批判性重新评估的必要性,并为改进数据来源与使用提供了方向。