Decoupling representation learning and classifier learning has been shown to be effective in classification with long-tailed data. There are two main ingredients in constructing a decoupled learning scheme; 1) how to train the feature extractor for representation learning so that it provides generalizable representations and 2) how to re-train the classifier that constructs proper decision boundaries by handling class imbalances in long-tailed data. In this work, we first apply Stochastic Weight Averaging (SWA), an optimization technique for improving the generalization of deep neural networks, to obtain better generalizing feature extractors for long-tailed classification. We then propose a novel classifier re-training algorithm based on stochastic representation obtained from the SWA-Gaussian, a Gaussian perturbed SWA, and a self-distillation strategy that can harness the diverse stochastic representations based on uncertainty estimates to build more robust classifiers. Extensive experiments on CIFAR10/100-LT, ImageNet-LT, and iNaturalist-2018 benchmarks show that our proposed method improves upon previous methods both in terms of prediction accuracy and uncertainty estimation.
翻译:将表示学习与分类器学习解耦已被证明在长尾数据分类中有效。构建解耦学习方案包含两个主要要素:1)如何训练用于表示学习的特征提取器,使其能够提供具有泛化能力的表示;2)如何通过处理长尾数据中的类别不平衡问题,重新训练能够构建恰当决策边界的分类器。在本工作中,我们首先应用随机权重平均(SWA)——一种提升深度神经网络泛化能力的优化技术——来获得更优泛化能力的特征提取器用于长尾分类。随后,我们提出一种新型分类器重训练算法,该算法基于从SWA-Gaussian(一种高斯扰动SWA)获得的随机表示,并采用自蒸馏策略,利用基于不确定性估计的多样化随机表示来构建更鲁棒的分类器。在CIFAR10/100-LT、ImageNet-LT和iNaturalist-2018基准上的大量实验表明,我们的方法在预测准确性和不确定性估计方面均优于此前方法。