Video transformers have recently emerged as a competitive alternative to 3D CNNs for video understanding. However, due to their large number of parameters and reduced inductive biases, these models require supervised pretraining on large-scale image datasets to achieve top performance. In this paper, we empirically demonstrate that self-supervised pretraining of video transformers on video-only datasets can lead to action recognition results that are on par or better than those obtained with supervised pretraining on large-scale image datasets, even massive ones such as ImageNet-21K. Since transformer-based models are effective at capturing dependencies over extended temporal spans, we propose a simple learning procedure that forces the model to match a long-term view to a short-term view of the same video. Our approach, named Long-Short Temporal Contrastive Learning (LSTCL), enables video transformers to learn an effective clip-level representation by predicting temporal context captured from a longer temporal extent. To demonstrate the generality of our findings, we implement and validate our approach under three different self-supervised contrastive learning frameworks (MoCo v3, BYOL, SimSiam) using two distinct video-transformer architectures, including an improved variant of the Swin Transformer augmented with space-time attention. We conduct a thorough ablation study and show that LSTCL achieves competitive performance on multiple video benchmarks and represents a convincing alternative to supervised image-based pretraining.


翻译:视频变压器最近成为了3DCNN视频理解的竞争性替代方案,但由于其参数数量众多,诱导偏差减少,这些模型需要接受大规模图像数据集的监督前培训,才能达到顶级性能。在本文中,我们从经验上表明,对视频变压器进行仅视频数据集的自我监督前培训,可以导致与大型图像数据集监督前培训相比的承认行动结果平等或更好,甚至像图像网21K这样的大型图像。由于基于变压器的模型能够有效地捕捉较长时间跨度的依赖性,因此我们建议了一个简单的学习程序,将模型与同一视频的短期视图相匹配。我们的方法,名为“长期光学变压变压式学习”(LSTLCL),使视频变压器能够通过预测从更长的时间范围所捕捉到的时间环境背景来学习有效的短期代表。为了显示我们发现的一般性,我们根据三种不同的自我监督反差对比学习框架(MOVV3)实施和验证我们的方法,我们提出一个简单的学习程序程序,一个不同的视频变压变压模型,一个不同的系统前SimStravial-travial Stal Stal Stal Stal Stal Stal Stal Stal Stal imal res一个不同的图像,包括一个不同的系统,一个不同的图像结构,一个清晰的Silvical-traview Staview

0
下载
关闭预览

相关内容

专知会员服务
90+阅读 · 2021年6月29日
【MIT】反偏差对比学习,Debiased Contrastive Learning
专知会员服务
92+阅读 · 2020年7月4日
【ICML2020】多视角对比图表示学习,Contrastive Multi-View GRL
专知会员服务
80+阅读 · 2020年6月11日
【Google】监督对比学习,Supervised Contrastive Learning
专知会员服务
75+阅读 · 2020年4月24日
Stabilizing Transformers for Reinforcement Learning
专知会员服务
61+阅读 · 2019年10月17日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
已删除
将门创投
7+阅读 · 2018年11月5日
Hierarchical Imitation - Reinforcement Learning
CreateAMind
19+阅读 · 2018年5月25日
Hierarchical Disentangled Representations
CreateAMind
4+阅读 · 2018年4月15日
Arxiv
7+阅读 · 2020年10月9日
Video-to-Video Synthesis
Arxiv
9+阅读 · 2018年8月20日
VIP会员
最新内容
刚刚!Jev中文教程项目发布了
专知会员服务
0+阅读 · 10月4日
《人工智能赋能的适应性多功能电磁战》
专知会员服务
12+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
11+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
17+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
相关资讯
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
已删除
将门创投
7+阅读 · 2018年11月5日
Hierarchical Imitation - Reinforcement Learning
CreateAMind
19+阅读 · 2018年5月25日
Hierarchical Disentangled Representations
CreateAMind
4+阅读 · 2018年4月15日
Top
微信扫码咨询专知VIP会员