Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural language and image understanding. We show that we can achieve significant improvements on audiovisual downstream classification tasks, surpassing the state-of-the-art on VGGSound and AudioSet. Furthermore, we can leverage our audiovisual pretraining scheme for multiple unimodal downstream tasks using a single audiovisual pretrained model. We additionally demonstrate the transferability of our representations, achieving state-of-the-art audiovisual results on Epic Kitchens without pretraining specifically for this dataset.
翻译:我们能否利用视频中已有的视听信息来改进自监督表征学习?为了回答这一问题,受自然语言和图像理解领域类似方法的成功启发,我们在掩码自编码框架下研究了多种预训练架构与目标。研究表明,我们能在视听下游分类任务上取得显著改进,在VGGSound和AudioSet上超越现有最佳水平。此外,利用单一视听预训练模型,我们可将该预训练方案应用于多个单模态下游任务。我们还展示了表征的可迁移性,在未针对Epic Kitchens数据集进行特定预训练的情况下,仍在该数据集上实现了最优的视听任务结果。