Data quality is crucial for training accurate, unbiased, and trustworthy machine learning models as well as for their correct evaluation. Recent works, however, have shown that even popular datasets used to train and evaluate state-of-the-art models contain a non-negligible amount of erroneous annotations, biases, or artifacts. While practices and guidelines regarding dataset creation projects exist, to our knowledge, large-scale analysis has yet to be performed on how quality management is conducted when creating natural language datasets and whether these recommendations are followed. Therefore, we first survey and summarize recommended quality management practices for dataset creation as described in the literature and provide suggestions for applying them. Then, we compile a corpus of 591 scientific publications introducing text datasets and annotate it for quality-related aspects, such as annotator management, agreement, adjudication, or data validation. Using these annotations, we then analyze how quality management is conducted in practice. A majority of the annotated publications apply good or excellent quality management. However, we deem the effort of 30\% of the works as only subpar. Our analysis also shows common errors, especially when using inter-annotator agreement and computing annotation error rates.
翻译:数据质量对于训练准确、无偏且可信的机器学习模型及其正确评估至关重要。然而,近期研究表明,即使用于训练和评估最先进模型的流行数据集也包含不可忽视的错误标注、偏差或伪影。尽管存在关于数据集创建项目的实践指南,但据我们所知,目前尚未有大规模分析研究在自然语言数据集创建过程中如何实施质量管理,以及这些建议是否得到遵循。为此,我们首先梳理并总结了文献中推荐的数据集创建质量管理实践,提出应用建议。随后,我们编制了包含591篇引入文本数据集的科学文献语料库,并对其标注质量相关要素(如标注人员管理、一致性评估、裁决机制或数据验证)。基于这些标注,我们分析了实际质量管理实施情况。大多数被标注文献采用了良好或卓越的质量管理措施,但30%的工作被认为质量仅属一般。分析还揭示了常见错误,特别是在使用标注者间一致性和计算标注错误率方面。