Transformer-based architectures have become competitive across a variety of visual domains, most notably images and videos. While prior work studies these modalities in isolation, having a common architecture suggests that one can train a single unified model for multiple visual modalities. Prior attempts at unified modeling typically use architectures tailored for vision tasks, or obtain worse performance compared to single modality models. In this work, we show that masked autoencoding can be used to train a simple Vision Transformer on images and videos, without requiring any labeled data. This single model learns visual representations that are comparable to or better than single-modality representations on both image and video benchmarks, while using a much simpler architecture. Furthermore, this model can be learned by dropping 90% of the image and 95% of the video patches, enabling extremely fast training of huge model architectures. In particular, we show that our single ViT-Huge model can be finetuned to achieve 86.6% on ImageNet and 75.5% on the challenging Something Something-v2 video benchmark, setting a new state-of-the-art.
翻译:基于Transformer的架构已在多种视觉领域(尤其是图像与视频)展现出竞争力。虽然先前研究通常独立处理这些模态,但共享的架构暗示可以针对多种视觉模态训练单一统一模型。早期的统一建模尝试通常采用专为视觉任务定制的架构,或相比单模态模型表现更差。本研究表明,掩码自编码可用于在图像和视频上训练简单的视觉Transformer(Vision Transformer),无需任何标注数据。该单一模型学到的视觉表示在图像和视频基准测试中可媲美甚至超越单模态表示,同时采用更简洁的架构。此外,该模型可仅通过丢弃90%的图像块和95%的视频块进行学习,从而极大加速超大规模架构的训练。具体而言,我们展示单个ViT-Huge模型微调后在ImageNet上达到86.6%的准确率,在极具挑战性的Something Something-v2视频基准上达到75.5%,刷新了当前最优水平。