Self-supervised learning is an effective way for label-free model pre-training, especially in the video domain where labeling is expensive. Existing self-supervised works in the video domain use varying experimental setups to demonstrate their effectiveness and comparison across approaches becomes challenging with no standard benchmark. In this work, we first provide a benchmark that enables a comparison of existing approaches on the same ground. Next, we study five different aspects of self-supervised learning important for videos; 1) dataset size, 2) complexity, 3) data distribution, 4) data noise, and, 5)feature analysis. To facilitate this study, we focus on seven different methods along with seven different network architectures and perform an extensive set of experiments on 5 different datasets with an evaluation of two different downstream tasks. We present several interesting insights from this study which span across different properties of pretraining and target datasets, pretext-tasks, and model architectures among others. We further put some of these insights to the real test and propose an approach that requires a limited amount of training data and outperforms existing state-of-the-art approaches which use 10x pretraining data. We believe this work will pave the way for researchers to a better understanding of self-supervised pretext tasks in video representation learning.
翻译:自监督学习是一种有效的无标签模型预训练方法,尤其在标注成本高昂的视频领域具有重要价值。现有视频领域的自监督工作采用各异的实验设置来验证其有效性,由于缺乏统一基准,不同方法间的比较变得困难。本研究首先构建了一个基准框架,使现有方法能在同一基础上进行比较。其次,我们系统研究了视频自监督学习中五个关键方面:1)数据集规模、2)复杂度、3)数据分布、4)数据噪声以及5)特征分析。为此,我们聚焦七种不同方法,配合七种网络架构,在五个数据集上开展大量实验,并通过两个下游任务进行评估。研究揭示了预训练与目标数据集特性、前置任务设计以及模型架构等多方面的有趣发现。我们进一步将部分发现付诸实践,提出了一种仅需有限训练数据的方法,其性能超越了使用10倍预训练数据的现有最优方法。本工作将为研究人员深入理解视频表示学习中的自监督前置任务奠定基础。