Self-Supervised learning (SSL) with Joint-Embedding Architectures (JEA) has led to outstanding performances. All instantiations of this paradigm were trained using strong and well-established hand-crafted data augmentations, leading to the general belief that they are required for the proper training and performance of such models. On the other hand, generative reconstruction-based models such as BEIT and MAE or Joint-Embedding Predictive Architectures such as I-JEPA have shown strong performance without using data augmentations except masking. In this work, we challenge the importance of invariance and data-augmentation in JEAs at scale. By running a case-study on a recent SSL foundation model - DINOv2 - we show that strong image representations can be obtained with JEAs and only cropping without resizing provided the training data is large enough, reaching state-of-the-art results and using the least amount of augmentation in the literature. Through this study, we also discuss the impact of compute constraints on the outcomes of experimental deep learning research, showing that they can lead to very different conclusions.
翻译:自监督学习(SSL)结合联合嵌入架构(JEA)已取得了卓越的性能。该范式的所有实例均使用强大且成熟的手工数据增强方法进行训练,导致普遍认为这些增强对于此类模型的正确训练和性能表现是必需的。另一方面,基于生成重建的模型(如BEIT和MAE)或联合嵌入预测架构(如I-JEPA)在不使用除掩码以外的数据增强的情况下也展现出了强劲性能。在本工作中,我们质疑了在规模化JEA中不变性和数据增强的重要性。通过以近期自监督学习基础模型——DINOv2——作为案例研究,我们证明:在训练数据足够大的情况下,仅使用不改变尺寸的裁剪操作,JEA即可获得强大的图像表示,达到文献中最先进的结果,同时使用了最少的增强量。通过这项研究,我们还讨论了计算资源限制对实验性深度学习研究结论的影响,表明其可能导致截然不同的结论。