Most Artificial Intelligence applications are based on supervised machine learning (ML), which ultimately grounds on manually annotated data. The annotation process is often performed in terms of a majority vote and this has been proved to be often problematic, as highlighted by recent studies on the evaluation of ML models. In this article we describe and advocate for a different paradigm, which we call data perspectivism, which moves away from traditional gold standard datasets, towards the adoption of methods that integrate the opinions and perspectives of the human subjects involved in the knowledge representation step of ML processes. Drawing on previous works which inspired our proposal we describe the potential of our proposal for not only the more subjective tasks (e.g. those related to human language) but also to tasks commonly understood as objective (e.g. medical decision making), and present the main advantages of adopting a perspectivist stance in ML, as well as possible disadvantages, and various ways in which such a stance can be implemented in practice. Finally, we share a set of recommendations and outline a research agenda to advance the perspectivist stance in ML.
翻译:大多数人工智能应用基于监督式机器学习(ML),而监督学习最终依赖于人工标注的数据。标注过程通常通过多数投票原则进行,但近期关于机器学习模型评估的研究表明,这种做法往往存在问题。本文描述并倡导一种不同的范式,我们称之为“数据视角主义”(data perspectivism),该范式摒弃传统的黄金标准数据集,转向采用能够整合机器学习知识表示步骤中人类主体观点与视角的方法。基于启发我们提出这一方案的前期工作,我们阐述了该方法的潜力——不仅适用于更为主观的任务(如与人类语言相关的任务),也适用于通常被视为客观的任务(如医疗决策),并论述了在机器学习中采纳视角主义立场的主要优势、潜在弊端,以及实践中实施该立场的多种途径。最后,我们提出一系列建议,并概述推进机器学习视角主义立场的研究议程。