Hyperbolic manifolds for visual representation learning allow for effective learning of semantic class hierarchies by naturally embedding tree-like structures with low distortion within a low-dimensional representation space. The highly separable semantic class hierarchies produced by hyperbolic learning have shown to be powerful in low-shot tasks, however, their application in self-supervised learning is yet to be explored fully. In this work, we explore the use of hyperbolic representation space for self-supervised representation learning for prototype-based clustering approaches. First, we extend the Masked Siamese Networks to operate on the Poincar\'e ball model of hyperbolic space, secondly, we place prototypes on the ideal boundary of the Poincar\'e ball. Unlike previous methods we project to the hyperbolic space at the output of the encoder network and utilise a hyperbolic projection head to ensure that the representations used for downstream tasks remain hyperbolic. Empirically we demonstrate the ability of these methods to perform comparatively to Euclidean methods in lower dimensions for linear evaluation tasks, whilst showing improvements in extreme few-shot learning tasks.
翻译:双曲流形用于视觉表征学习,可通过在低维表示空间中以低失真自然嵌入树状结构,有效学习语义类别层次结构。双曲学习产生的高度可分离语义类别层次结构在低样本任务中表现出强大能力,然而其在自监督学习中的应用仍有待充分探索。本文探索了双曲表示空间用于基于原型聚类的自监督表示学习。首先,我们将掩码孪生网络扩展至双曲空间的庞加莱球模型运作;其次,将原型置于庞加莱球的理想边界上。与先前方法不同,我们在编码器网络输出端向双曲空间投影,并利用双曲投影头确保下游任务使用的表示保持双曲特性。实验表明,该方法在线性评估任务中能以更低维度与欧几里得方法性能相当,同时在极端少样本学习任务中展现出显著改进。