Partial-label learning (PLL) relies on a key assumption that the true label of each training example must be in the candidate label set. This restrictive assumption may be violated in complex real-world scenarios, and thus the true label of some collected examples could be unexpectedly outside the assigned candidate label set. In this paper, we term the examples whose true label is outside the candidate label set OOC (out-of-candidate) examples, and pioneer a new PLL study to learn with OOC examples. We consider two types of OOC examples in reality, i.e., the closed-set/open-set OOC examples whose true label is inside/outside the known label space. To solve this new PLL problem, we first calculate the wooden cross-entropy loss from candidate and non-candidate labels respectively, and dynamically differentiate the two types of OOC examples based on specially designed criteria. Then, for closed-set OOC examples, we conduct reversed label disambiguation in the non-candidate label set; for open-set OOC examples, we leverage them for training by utilizing an effective regularization strategy that dynamically assigns random candidate labels from the candidate label set. In this way, the two types of OOC examples can be differentiated and further leveraged for model training. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art PLL methods.
翻译:偏标签学习(PLL)依赖于一个关键假设:每个训练样本的真实标签必须位于候选标签集中。这一限制性假设在复杂真实场景中可能被违背,因此部分收集样本的真实标签可能意外地不在分配的候选标签集中。本文将真实标签位于候选标签集之外的样本称为候选外(OOC)示例,并开创性地研究如何利用OOC示例进行PLL学习。我们考虑现实中两类OOC示例:真实标签位于已知标签空间内的闭集OOC示例,以及真实标签位于已知标签空间外的开集OOC示例。为解决这一新型PLL问题,我们首先分别计算候选标签和非候选标签的木交叉熵损失,并基于专门设计的准则动态区分两类OOC示例。随后,针对闭集OOC示例,我们在非候选标签集中进行反向标签消歧;针对开集OOC示例,我们采用有效的正则化策略动态分配候选标签集中的随机候选标签,将其用于模型训练。通过这种方式,两类OOC示例可被区分并进一步用于模型训练。大量实验表明,我们提出的方法优于现有最优的PLL方法。