The negative impact of label noise is well studied in classical supervised learning yet remains an open research question in meta-learning. Meta-learners aim to adapt to unseen learning tasks by learning a good initial model in meta-training and consecutively fine-tuning it according to new tasks during meta-testing. In this paper, we present the first extensive analysis of the impact of varying levels of label noise on the performance of state-of-the-art meta-learners, specifically gradient-based $N$-way $K$-shot learners. We show that the accuracy of Reptile, iMAML, and foMAML drops by up to 42% on the Omniglot and CifarFS datasets when meta-training is affected by label noise. To strengthen the resilience against label noise, we propose two sampling techniques, namely manifold (Man) and batch manifold (BatMan), which transform the noisy supervised learners into semi-supervised ones to increase the utility of noisy labels. We first construct manifold samples of $N$-way $2$-contrastive-shot tasks through augmentation, learning the embedding via a contrastive loss in meta-training, and then perform classification through zeroing on the embedding in meta-testing. We show that our approach can effectively mitigate the impact of meta-training label noise. Even with 60% wrong labels \batman and \man can limit the meta-testing accuracy drop to ${2.5}$, ${9.4}$, ${1.1}$ percent points, respectively, with existing meta-learners across the Omniglot, CifarFS, and MiniImagenet datasets.
翻译:摘要:标签噪声的负面影响在经典监督学习中已得到充分研究,但在元学习中仍是一个待解决的研究问题。元学习器旨在通过元训练阶段学习良好的初始模型,随后在元测试阶段根据新任务对其进行微调,从而适应未见过的学习任务。本文首次深入分析了不同级别的标签噪声对最先进元学习器性能的影响,特别是基于梯度的$N$路$K$样本学习方法。我们证明,当元训练受到标签噪声影响时,Reptile、iMAML和foMAML在Omniglot和CifarFS数据集上的准确率下降高达42%。为增强对标签噪声的鲁棒性,我们提出两种采样技术,即流形采样(Man)和批流形采样(BatMan),将有噪声的监督学习器转化为半监督学习器,以提高噪声标签的效用。我们首先通过数据增强构建$N$路$2$对比样本任务的流形样本,在元训练中通过对比损失学习嵌入,随后在元测试中通过嵌入置零进行分类。我们证明,该方法能有效缓解元训练标签噪声的影响。即使标签错误率达到60%,BatMan和Man方法可将现有元学习器在Omniglot、CifarFS和MiniImagenet数据集上的元测试准确率下降分别限制在2.5、9.4、1.1个百分点以内。