Learning from crowds describes that the annotations of training data are obtained with crowd-sourcing services. Multiple annotators each complete their own small part of the annotations, where labeling mistakes that depend on annotators occur frequently. Modeling the label-noise generation process by the noise transition matrix is a power tool to tackle the label noise. In real-world crowd-sourcing scenarios, noise transition matrices are both annotator- and instance-dependent. However, due to the high complexity of annotator- and instance-dependent transition matrices (AIDTM), annotation sparsity, which means each annotator only labels a little part of instances, makes modeling AIDTM very challenging. Prior works simplify the problem by assuming the transition matrix is instance-independent or using simple parametric ways, which lose modeling generality. Motivated by this, we target a more realistic problem, estimating general AIDTM in practice. Without losing modeling generality, we parameterize AIDTM with deep neural networks. To alleviate the modeling challenge, we suppose every annotator shares its noise pattern with similar annotators, and estimate AIDTM via knowledge transfer. We hence first model the mixture of noise patterns by all annotators, and then transfer this modeling to individual annotators. Furthermore, considering that the transfer from the mixture of noise patterns to individuals may cause two annotators with highly different noise generations to perturb each other, we employ the knowledge transfer between identified neighboring annotators to calibrate the modeling. Theoretical analyses are derived to demonstrate that both the knowledge transfer from global to individuals and the knowledge transfer between neighboring individuals can help model general AIDTM. Experiments confirm the superiority of the proposed approach on synthetic and real-world crowd-sourcing data.
翻译:从众包学习描述了训练数据的标注是通过众包服务获得的。多个标注者各自完成其小部分标注任务,其中依赖标注者的标注错误频繁发生。利用噪声转移矩阵对标签噪声生成过程进行建模是处理标签噪声的有力工具。在真实众包场景中,噪声转移矩阵同时依赖于标注者和实例。然而,由于标注者与实例相关转移矩阵(AIDTM)的高度复杂性,标注稀疏性(即每个标注者仅标注少量实例)使得对AIDTM的建模极具挑战性。以往的工作通过假设转移矩阵与实例无关或采用简单的参数化方法简化问题,但损失了建模的通用性。受此启发,我们针对一个更实际的问题,即估计实践中通用的AIDTM。在不损失建模通用性的前提下,我们利用深度神经网络对AIDTM进行参数化。为缓解建模挑战,我们假设每个标注者与相似标注者共享噪声模式,并通过知识迁移估计AIDTM。因此,我们首先对所有标注者的噪声模式混合进行建模,然后将此建模迁移至单个标注者。此外,考虑到从噪声模式混合到个体的迁移可能导致噪声生成差异较大的两个标注者相互干扰,我们利用已识别的相邻标注者之间的知识迁移来校准建模。理论分析表明,从全局到个体的知识迁移以及相邻个体之间的知识迁移均有助于建模通用的AIDTM。实验在合成和真实众包数据上验证了所提方法的优越性。