Label noise is a common issue in real-world datasets that inevitably impacts the generalization of models. This study focuses on robust classification tasks where the label noise is instance-dependent. Estimating the transition matrix accurately in this task is challenging, and methods based on sample selection often exhibit confirmation bias to varying degrees. Sparse over-parameterized training (SOP) has been theoretically effective in estimating and recovering label noise, offering a novel solution for noise-label learning. However, this study empirically observes and verifies a technical flaw of SOP: the lack of coordination between model predictions and noise recovery leads to increased generalization error. To address this, we propose a method called Coordinated Sparse Recovery (CSR). CSR introduces a collaboration matrix and confidence weights to coordinate model predictions and noise recovery, reducing error leakage. Based on CSR, this study designs a joint sample selection strategy and constructs a comprehensive and powerful learning framework called CSR+. CSR+ significantly reduces confirmation bias, especially for datasets with more classes and a high proportion of instance-specific noise. Experimental results on simulated and real-world noisy datasets demonstrate that both CSR and CSR+ achieve outstanding performance compared to methods at the same level.
翻译:标签噪声是现实世界数据集中常见的问题,不可避免地会影响模型的泛化能力。本研究聚焦于标签噪声与实例相关的鲁棒分类任务。在此任务中,准确估计转移矩阵具有挑战性,而基于样本选择的方法往往存在不同程度的确认偏差。稀疏过参数化训练(SOP)在理论上能有效估计和恢复标签噪声,为噪声标签学习提供了新颖的解决方案。然而,本研究发现并验证了SOP的一个技术缺陷:模型预测与噪声恢复之间缺乏协调,导致泛化误差增大。为此,我们提出了一种名为协同稀疏恢复(CSR)的方法。CSR引入协作矩阵和置信权重,用以协调模型预测与噪声恢复,从而减少误差传播。基于CSR,本研究设计了一种联合样本选择策略,并构建了一个全面且强大的学习框架CSR+。CSR+显著降低了确认偏差,尤其适用于类别数量较多且实例特定噪声比例较高的数据集。在模拟和真实噪声数据集上的实验结果表明,CSR和CSR+相比同级别方法均取得了卓越性能。