In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and spatial representations from unlabeled spatial audios, thereby enhancing both event classification and sound localization in downstream tasks. At its core, we propose a multi-level data augmentation pipeline that augments different levels of audio features, including waveforms, Mel spectrograms, and generalized cross-correlation (GCC) features. In addition, we introduce simple yet effective channel-wise augmentation methods to randomly swap the order of the microphones and mask Mel and GCC channels. By using these augmentations, we find that linear layers on top of the learned representation significantly outperform supervised models in terms of both event classification accuracy and localization error. We also perform a comprehensive analysis of the effect of each augmentation method and a comparison of the fine-tuning performance using different amounts of labeled data.
翻译:在本研究中,我们提出了一种用于对比学习的简单多通道框架(MC-SimCLR),以编码空间音频的“是什么”和“在哪里”。MC-SimCLR从未标记的空间音频中学习联合频谱和空间表示,从而在下游任务中同时提升事件分类和声音定位的性能。其核心在于,我们提出了一种多级数据增强流水线,对音频特征的不同层次(包括波形、梅尔频谱图和广义互相关特征)进行增强。此外,我们引入了简单而有效的通道级增强方法,用于随机交换麦克风顺序并掩码梅尔通道和GCC通道。通过使用这些增强方法,我们发现,基于所学表示的线性层在事件分类准确率和定位误差方面均显著优于有监督模型。我们还对每种增强方法的效果进行了全面分析,并比较了使用不同数量标记数据时的微调性能。