We introduce a regularization loss based on kernel mean embeddings with rotation-invariant kernels on the hypersphere (also known as dot-product kernels) for self-supervised learning of image representations. Besides being fully competitive with the state of the art, our method significantly reduces time and memory complexity for self-supervised training, making it implementable for very large embedding dimensions on existing devices and more easily adjustable than previous methods to settings with limited resources. Our work follows the major paradigm where the model learns to be invariant to some predefined image transformations (cropping, blurring, color jittering, etc.), while avoiding a degenerate solution by regularizing the embedding distribution. Our particular contribution is to propose a loss family promoting the embedding distribution to be close to the uniform distribution on the hypersphere, with respect to the maximum mean discrepancy pseudometric. We demonstrate that this family encompasses several regularizers of former methods, including uniformity-based and information-maximization methods, which are variants of our flexible regularization loss with different kernels. Beyond its practical consequences for state-of-the-art self-supervised learning with limited resources, the proposed generic regularization approach opens perspectives to leverage more widely the literature on kernel methods in order to improve self-supervised learning methods.
翻译:我们提出了一种基于超球面上旋转不变核(也称为点积核)的核均值嵌入正则化损失函数,用于图像表示的自监督学习。该方法不仅与现有最优方法完全竞争,还显著降低了自监督训练的时间和内存复杂度,使其能够在现有设备上处理极大嵌入维度,并且比先前方法更易于适应资源受限的场景。我们的工作遵循主流范式:模型学会对某些预定义的图像变换(裁剪、模糊、颜色抖动等)保持不变性,同时通过正则化嵌入分布来避免退化解。我们的核心贡献是提出一种损失函数族,该函数族基于最大均值差异伪度量,促使嵌入分布逼近超球面上的均匀分布。我们证明该函数族涵盖了先前方法的多种正则化器,包括基于均匀性和基于信息最大化的方法——这些方法实为不同核函数下我们所提柔性正则化损失的变体。除对资源受限场景下最优自监督学习具有实际意义外,所提出的通用正则化方法还为更广泛地利用核方法文献以改进自监督学习开辟了新视角。