Despite impressive empirical advances of SSL in solving various tasks, the problem of understanding and characterizing SSL representations learned from input data remains relatively under-explored. We provide a comparative analysis of how the representations produced by SSL models differ when masking parts of the input. Specifically, we considered state-of-the-art SSL pretrained models, such as DINOv2, MAE, and SwaV, and analyzed changes at the representation levels across 4 Image Classification datasets. First, we generate variations of the datasets by applying foreground and background segmentation. Then, we conduct statistical analysis using Canonical Correlation Analysis (CCA) and Centered Kernel Alignment (CKA) to evaluate the robustness of the representations learned in SSL models. Empirically, we show that not all models lead to representations that separate foreground, background, and complete images. Furthermore, we test different masking strategies by occluding the center regions of the images to address cases where foreground and background are difficult. For example, the DTD dataset that focuses on texture rather specific objects.
翻译:尽管SSL在解决各种任务方面取得了令人瞩目的实证进展,但理解和表征从输入数据中学习到的SSL表示的问题仍然相对未被充分探索。我们提供了一项比较分析,研究当输入部分被遮蔽时,SSL模型产生的表示如何变化。具体来说,我们考虑了最先进的SSL预训练模型,如DINOv2、MAE和SwaV,并在4个图像分类数据集上分析了表示层面的变化。首先,通过应用前景和背景分割,我们生成了数据集的变体。然后,我们使用典型相关分析(CCA)和中心核对齐(CKA)进行统计分析,以评估SSL模型中学习到的表示的鲁棒性。通过实证研究,我们表明并非所有模型都能产生分离前景、背景和完整图像的表示。此外,我们测试了不同的遮蔽策略,通过遮挡图像的中心区域来处理前景和背景难以区分的情况。例如,专注于纹理而非特定对象的DTD数据集。