In recent years, ML researchers have wrestled with defining and improving machine learning (ML) benchmarks and datasets. In parallel, some have trained a critical lens on the ethics of dataset creation and ML research. In this position paper, we highlight the entanglement of ethics with seemingly ``technical'' or ``scientific'' decisions about the design of ML benchmarks. Our starting point is the existence of multiple overlooked structural similarities between human intelligence benchmarks and ML benchmarks. Both types of benchmarks set standards for describing, evaluating, and comparing performance on tasks relevant to intelligence -- standards that many scholars of human intelligence have long recognized as value-laden. We use perspectives from feminist philosophy of science on IQ benchmarks and thick concepts in social science to argue that values need to be considered and documented when creating ML benchmarks. It is neither possible nor desirable to avoid this choice by creating value-neutral benchmarks. Finally, we outline practical recommendations for ML benchmark research ethics and ethics review.
翻译:近年来,机器学习研究人员致力于定义和改进机器学习基准与数据集。与此同时,部分研究者开始以批判性视角审视数据集创建和机器学习研究中的伦理问题。在本观点论文中,我们着重揭示伦理问题如何与看似"技术性"或"科学性"的机器学习基准设计决策相互交织。我们的出发点在于:人类智能基准与机器学习基准之间存在着若干被长期忽视的结构相似性。两类基准均设定了描述、评估和比较与智能相关任务表现的标准——众多人类智能研究者早已认识到这些标准蕴含价值判断。我们借鉴女性主义科学哲学关于智商基准的研究视角,以及社会科学中的"厚概念"理论,论证在创建机器学习基准时必须考虑并记录价值因素。试图通过构建价值中立的基准来回避这一选择,既不可行亦不可取。最后,我们为机器学习基准研究的伦理规范与伦理审查提出实践性建议。