Masked AutoEncoder (MAE) has revolutionized the field of self-supervised learning with its simple yet effective masking and reconstruction strategies. However, despite achieving state-of-the-art performance across various downstream vision tasks, the underlying mechanisms that drive MAE's efficacy are less well-explored compared to the canonical contrastive learning paradigm. In this paper, we first propose a local perspective to explicitly extract a local contrastive form from MAE's reconstructive objective at the patch level. And then we introduce a new empirical framework, called Local Contrastive MAE (LC-MAE), to analyze both reconstructive and contrastive aspects of MAE. LC-MAE reveals that MAE learns invariance to random masking and ensures distribution consistency between the learned token embeddings and the original images. Furthermore, we dissect the contribution of the decoder and random masking to MAE's success, revealing both the decoder's learning mechanism and the dual role of random masking as data augmentation and effective receptive field restriction. Our experimental analysis sheds light on the intricacies of MAE and summarizes some useful design methodologies, which can inspire more powerful visual self-supervised methods.
翻译:掩码自编码器(MAE)凭借其简单而有效的掩码与重建策略,彻底革新了自监督学习领域。然而,尽管MAE在各类下游视觉任务中取得了最优性能,其作用机制相比经典的对比学习范式仍缺乏深入探索。本文首先提出一种局部视角,从MAE重建目标的补丁层级显式提取局部对比形式;进而引入名为局部对比MAE(LC-MAE)的新型实证框架以分析MAE的重建性与对比性双重特性。LC-MAE揭示:MAE能够学习对随机掩码的不变性,并确保学习到的令牌嵌入与原始图像间的分布一致性。此外,我们剖析了解码器与随机掩码对MAE成功的贡献机制,不仅阐明了解码器的学习机理,更揭示了随机掩码既充当数据增强又限制有效感受野的双重作用。本实验分析揭示了MAE的内在复杂性,并总结出若干实用设计方法论,可为开发更强大的视觉自监督方法提供启发。