Contrastive self-supervised learning has gained attention for its ability to create high-quality representations from large unlabelled data sets. A key reason that these powerful features enable data-efficient learning of downstream tasks is that they provide augmentation invariance, which is often a useful inductive bias. However, the amount and type of invariances preferred is not known apriori, and varies across different downstream tasks. We therefore propose a multi-task self-supervised framework (MT-SLVR) that learns both variant and invariant features in a parameter-efficient manner. Our multi-task representation provides a strong and flexible feature that benefits diverse downstream tasks. We evaluate our approach on few-shot classification tasks drawn from a variety of audio domains and demonstrate improved classification performance on all of them
翻译:对比自监督学习因其能够从大规模无标注数据中生成高质量表征而受到关注。这些强大特征之所以能够实现下游任务的高效数据学习,关键在于其提供了增强不变性——这一常用且有效的归纳偏置。然而,偏好的不变性类型与程度并非先验可知,且因下游任务而异。为此,我们提出了一种多任务自监督学习框架(MT-SLVR),以参数高效的方式同时学习变换可变与不变特征。这种多任务表征提供了兼具强健性与灵活性的特征,可满足多样化下游任务的需求。我们在多个音频领域的少样本分类任务上评估了所提方法,结果表明该方法在所有任务中均提升了分类性能。