Self-supervised learning (SSL) is a powerful tool in machine learning, but understanding the learned representations and their underlying mechanisms remains a challenge. This paper presents an in-depth empirical analysis of SSL-trained representations, encompassing diverse models, architectures, and hyperparameters. Our study reveals an intriguing aspect of the SSL training process: it inherently facilitates the clustering of samples with respect to semantic labels, which is surprisingly driven by the SSL objective's regularization term. This clustering process not only enhances downstream classification but also compresses the data information. Furthermore, we establish that SSL-trained representations align more closely with semantic classes rather than random classes. Remarkably, we show that learned representations align with semantic classes across various hierarchical levels, and this alignment increases during training and when moving deeper into the network. Our findings provide valuable insights into SSL's representation learning mechanisms and their impact on performance across different sets of classes.
翻译:自监督学习(SSL)是机器学习中一种强大的工具,但理解学习到的表征及其内在机制仍是一项挑战。本文对SSL训练的表征进行了深入的实证分析,涵盖不同的模型、架构和超参数。我们的研究揭示了SSL训练过程中一个引人注目的方面:它本质上有助于根据语义标签对样本进行聚类,这一现象令人惊讶地由SSL目标的正则化项驱动。该聚类过程不仅提升了下游分类性能,还压缩了数据信息。此外,我们确认SSL训练的表征与语义类别(而非随机类别)更为一致。值得注意的是,我们证明学习到的表征在不同层级上均与语义类别对齐,并且这种对齐度会随着训练过程及网络层数的加深而增强。我们的发现为SSL的表征学习机制及其对不同类别集合性能的影响提供了重要见解。