Learning from noisy data has attracted much attention, where most methods focus on closed-set label noise. However, a more common scenario in the real world is the presence of both open-set and closed-set noise. Existing methods typically identify and handle these two types of label noise separately by designing a specific strategy for each type. However, in many real-world scenarios, it would be challenging to identify open-set examples, especially when the dataset has been severely corrupted. Unlike the previous works, we explore how models behave when faced with open-set examples, and find that \emph{a part of open-set examples gradually get integrated into certain known classes}, which is beneficial for the separation among known classes. Motivated by the phenomenon, we propose a novel two-step contrastive learning method CECL (Class Expansion Contrastive Learning) which aims to deal with both types of label noise by exploiting the useful information of open-set examples. Specifically, we incorporate some open-set examples into closed-set classes to enhance performance while treating others as delimiters to improve representative ability. Extensive experiments on synthetic and real-world datasets with diverse label noise demonstrate the effectiveness of CECL.
翻译:从噪声数据中学习已引起广泛关注,大多数方法聚焦于封闭集标签噪声。然而,现实世界中更常见的场景是同时存在开放集噪声和封闭集噪声。现有方法通常通过为每种类型设计特定策略,分别识别和处理这两类标签噪声。但在许多现实场景中,识别开放集样本具有挑战性,尤其是当数据集受到严重污染时。不同于以往的工作,我们探究了模型在面对开放集样本时的行为,并发现部分开放集样本逐渐融入某些已知类别,这有利于已知类别之间的分离。受此现象启发,我们提出了一种新颖的两步对比学习方法CECL(类扩展对比学习),旨在通过利用开放集样本的有用信息来处理两类标签噪声。具体而言,我们将部分开放集样本融入封闭集类别以增强性能,同时将其他样本作为分隔符以提升表示能力。在合成和真实数据集上进行的广泛实验,涵盖了多种标签噪声类型,验证了CECL的有效性。