Using noisy crowdsourced labels from multiple annotators, a deep learning-based end-to-end (E2E) system aims to learn the label correction mechanism and the neural classifier simultaneously. To this end, many E2E systems concatenate the neural classifier with multiple annotator-specific ``label confusion'' layers and co-train the two parts in a parameter-coupled manner. The formulated coupled cross-entropy minimization (CCEM)-type criteria are intuitive and work well in practice. Nonetheless, theoretical understanding of the CCEM criterion has been limited. The contribution of this work is twofold: First, performance guarantees of the CCEM criterion are presented. Our analysis reveals for the first time that the CCEM can indeed correctly identify the annotators' confusion characteristics and the desired ``ground-truth'' neural classifier under realistic conditions, e.g., when only incomplete annotator labeling and finite samples are available. Second, based on the insights learned from our analysis, two regularized variants of the CCEM are proposed. The regularization terms provably enhance the identifiability of the target model parameters in various more challenging cases. A series of synthetic and real data experiments are presented to showcase the effectiveness of our approach.
翻译:利用来自多个标注者的含噪众包标签,基于深度学习的端到端(E2E)系统旨在同时学习标签纠正机制和神经分类器。为此,许多E2E系统将神经分类器与多个标注者特定的“标签混淆”层串联,并以参数耦合的方式共同训练这两部分。所构建的耦合交叉熵最小化(CCEM)型准则直观且在实践中表现良好。然而,对CCEM准则的理论理解一直有限。本文的贡献有两方面:首先,给出了CCEM准则的性能保证。我们的分析首次揭示,在现实条件下(例如,仅当不完整标注者标注和有限样本可用时),CCEM确实能正确识别标注者的混淆特征和所需的“真实标签”神经分类器。其次,基于分析所得洞察,提出了CCEM的两个正则化变体。这些正则化项可证明地在各种更具挑战性的情况下增强目标模型参数的可辨识性。通过一系列合成和真实数据实验,展示了我们方法的有效性。