Deep learning has been widely used recently for sound event detection and classification. Its success is linked to the availability of sufficiently large datasets, possibly with corresponding annotations when supervised learning is considered. In bioacoustic applications, most tasks come with few labelled training data, because annotating long recordings is time consuming and costly. Therefore supervised learning is not the best suited approach to solve bioacoustic tasks. The bioacoustic community recasted the problem of sound event detection within the framework of few-shot learning, i.e. training a system with only few labeled examples. The few-shot bioacoustic sound event detection task in the DCASE challenge focuses on detecting events in long audio recordings given only five annotated examples for each class of interest. In this paper, we show that learning a rich feature extractor from scratch can be achieved by leveraging data augmentation using a supervised contrastive learning framework. We highlight the ability of this framework to transfer well for five-shot event detection on previously unseen classes in the training data. We obtain an F-score of 63.46\% on the validation set and 42.7\% on the test set, ranking second in the DCASE challenge. We provide an ablation study for the critical choices of data augmentation techniques as well as for the learning strategy applied on the training set.
翻译:近年来,深度学习已广泛应用于声音事件检测与分类。其成功依赖于足够大规模数据集的可用性,且在采用监督学习时通常需要相应的标注信息。在生物声学应用中,由于对长时间录音进行标注耗时且成本高昂,大多数任务仅有少量标注训练数据。因此,监督学习并非解决生物声学任务的最优方法。生物声学界将声音事件检测问题重新定义为少样本学习框架下的任务,即仅利用少量标注样本训练系统。DCASE挑战中的少样本生物声学声音事件检测任务,要求在给定每个目标类别仅五个标注示例的条件下,从长音频录音中检测事件。本文证明,通过利用监督对比学习框架并结合数据增强,可以从零开始学习丰富的特征提取器。我们凸显了该框架对训练数据中未见类别进行五样本事件检测的良好迁移能力。在验证集上F-score达到63.46%,测试集上达到42.7%,在DCASE挑战中排名第二。我们还针对训练集上采用的数据增强技术关键选择及学习策略进行了消融研究。