Self-distillation methods using Siamese networks are popular for self-supervised pre-training. DINO is one such method based on a cross-entropy loss between $K$-dimensional probability vectors, obtained by applying a softmax function to the dot product between representations and learnt prototypes. Given the fact that the learned representations are $L^2$-normalized, we show that DINO and its derivatives, such as iBOT, can be interpreted as a mixture model of von Mises-Fisher components. With this interpretation, DINO assumes equal precision for all components when the prototypes are also $L^2$-normalized. Using this insight we propose DINO-vMF, that adds appropriate normalization constants when computing the cluster assignment probabilities. Unlike DINO, DINO-vMF is stable also for the larger ViT-Base model with unnormalized prototypes. We show that the added flexibility of the mixture model is beneficial in terms of better image representations. The DINO-vMF pre-trained model consistently performs better than DINO on a range of downstream tasks. We obtain similar improvements for iBOT-vMF vs iBOT and thereby show the relevance of our proposed modification also for other methods derived from DINO.
翻译:摘要:基于孪生网络的自蒸馏方法在自监督预训练中广受欢迎。DINO 是其中一种方法,它通过将表示与学习到的原型之间的点积应用 softmax 函数得到的 $K$ 维概率向量之间的交叉熵损失进行优化。鉴于学习到的表示经过 $L^2$ 归一化,我们证明 DINO 及其衍生方法(如 iBOT)可被解释为冯·米塞斯-费舍尔分量的混合模型。根据这一解释,当原型也经过 $L^2$ 归一化时,DINO 假设所有分量具有相同的精度。基于这一见解,我们提出了 DINO-vMF,它在计算聚类分配概率时加入了适当的归一化常数。与 DINO 不同,DINO-vMF 在使用未归一化原型的更大 ViT-Base 模型时也保持稳定。我们证明了混合模型的额外灵活性有助于获得更好的图像表示。在一系列下游任务中,DINO-vMF 预训练模型的性能始终优于 DINO。我们还观察到 iBOT-vMF 相对于 iBOT 的类似改进,从而表明我们提出的修改对于其他源自 DINO 的方法也具有相关性。