There has been much interest in recent years in learning good classifiers from data with noisy labels. Most work on learning from noisy labels has focused on standard loss-based performance measures. However, many machine learning problems require using non-decomposable performance measures which cannot be expressed as the expectation or sum of a loss on individual examples; these include for example the H-mean, Q-mean and G-mean in class imbalance settings, and the Micro $F_1$ in information retrieval. In this paper, we design algorithms to learn from noisy labels for two broad classes of multiclass non-decomposable performance measures, namely, monotonic convex and ratio-of-linear, which encompass all the above examples. Our work builds on the Frank-Wolfe and Bisection based methods of Narasimhan et al. (2015). In both cases, we develop noise-corrected versions of the algorithms under the widely studied family of class-conditional noise models. We provide regret (excess risk) bounds for our algorithms, establishing that even though they are trained on noisy data, they are Bayes consistent in the sense that their performance converges to the optimal performance w.r.t. the clean (non-noisy) distribution. Our experiments demonstrate the effectiveness of our algorithms in handling label noise.
翻译:近年来,从含噪标签数据中学习高质量分类器引起了广泛关注。大多数含噪标签学习研究聚焦于基于损失的标准性能度量。然而,许多机器学习问题需要使用非可分解性能度量——这类度量无法表示为单个样本损失的期望或求和,例如类别不平衡场景下的H-均值、Q-均值与G-均值,以及信息检索中的Micro $F_1$。本文针对两类广泛的多类非可分解性能度量(单调凸度量与线性比度量)设计含噪标签学习算法,上述示例均包含其中。我们的工作基于Narasimhan等人(2015)提出的Frank-Wolfe与二分法框架,并针对两类算法分别构建了基于广泛研究的类别条件噪声模型的噪声修正版本。我们给出了算法的遗憾(超额风险)界,证明尽管算法基于含噪数据训练,但其性能收敛于基于干净(无噪)分布的最优性能,即具有贝叶斯一致性。实验表明,本文算法在处理标签噪声方面具有显著有效性。