Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management and thereby reliable quality estimates are needed. Then, if quality is insufficient during the annotation process, rectifying measures can be taken to improve it. Quality estimation is often performed by having experts manually label instances as correct or incorrect. But checking all annotated instances tends to be expensive. Therefore, in practice, usually only subsets are inspected; sizes are chosen mostly without justification or regard to statistical power and more often than not, are relatively small. Basing estimates on small sample sizes, however, can lead to imprecise values for the error rate. Using unnecessarily large sample sizes costs money that could be better spent, for instance on more annotations. Therefore, we first describe in detail how to use confidence intervals for finding the minimal sample size needed to estimate the annotation error rate. Then, we propose applying acceptance sampling as an alternative to error rate estimation We show that acceptance sampling can reduce the required sample sizes up to 50% while providing the same statistical guarantees.
翻译:标注数据集是训练、评估、比较及部署监督式机器学习模型的关键要素。因此,确保标注的高质量至关重要。在创建过程中,需要良好的质量管理以及由此产生的可靠质量评估。若标注过程中质量不足,则可采取纠正措施加以改进。质量评估通常通过让专家手动将实例标记为正确或错误来完成。然而,检查所有已标注实例往往成本高昂。因此,实践中通常仅检查子集;样本大小的选择大多缺乏依据或未考虑统计功效,且往往相对较小。然而,基于小样本量得出的估计可能导致错误率数值不精确。使用非必要的大样本量则会耗费本可更有效利用的资金,例如用于更多标注。为此,我们首先详细描述如何利用置信区间确定估计标注错误率所需的最小样本量。接着,我们提出应用接受抽样作为错误率估计的替代方案。我们证明,接受抽样可在降低高达50%所需样本量的同时,提供相同的统计保证。