Human annotations are vital to supervised learning, yet annotators often disagree on the correct label, especially as annotation tasks increase in complexity. A strategy to improve label quality is to ask multiple annotators to label the same item and aggregate their labels. Many aggregation models have been proposed for categorical or numerical annotation tasks, but far less work has considered more complex annotation tasks involving open-ended, multivariate, or structured responses. While a variety of bespoke models have been proposed for specific tasks, our work is the first to introduce aggregation methods that generalize across many diverse complex tasks, including sequence labeling, translation, syntactic parsing, ranking, bounding boxes, and keypoints. This generality is achieved by devising a task-agnostic method to model distances between labels rather than the labels themselves. This article extends our prior work with investigation of three new research questions. First, how do complex annotation properties impact aggregation accuracy? Second, how should a task owner navigate the many modeling choices to maximize aggregation accuracy? Finally, what diagnoses can verify that aggregation models are specified correctly for the given data? To understand how various factors impact accuracy and to inform model selection, we conduct simulation studies and experiments on real, complex datasets. Regarding testing, we introduce unit tests for aggregation models and present a suite of such tests to ensure that a given model is not mis-specified and exhibits expected behavior. Beyond investigating these research questions above, we discuss the foundational concept of annotation complexity, present a new aggregation model as a bridge between traditional models and our own, and contribute a new semi-supervised learning method for complex label aggregation that outperforms prior work.
翻译:人工注释对监督学习至关重要,但注释者常因任务复杂度提升而对正确标签产生分歧。提高标签质量的策略之一是邀请多位注释者为同一项目标注并聚合其标签。已有诸多聚合模型针对分类或数值型注释任务提出,但针对涉及开放式、多变量或结构化响应的复杂注释任务的研究则显著不足。尽管针对特定任务存在多种定制模型,本研究首次提出能跨序列标注、翻译、句法分析、排序、边界框及关键点等多样复杂任务泛化的聚合方法。这种通用性通过设计任务无关的标签间距离建模方法(而非直接建模标签本身)实现。本文在前期工作基础上进一步探究三个新问题:其一,复杂注释属性如何影响聚合精度?其二,任务管理者应如何权衡多种建模选择以最大化聚合精度?其三,哪些诊断方法可验证聚合模型对给定数据的正确设定?为理解各因素对精度的影响并指导模型选择,我们开展仿真研究及真实复杂数据集实验。在模型验证方面,我们提出聚合模型的单元测试体系,通过一套标准化测试确保模型规范正确且行为符合预期。除上述研究问题外,本文还探讨注释复杂性的基础概念,提出作为传统模型与本文模型桥梁的新型聚合方法,并贡献一种优于现有方法的半监督复杂标签聚合学习技术。