Self-predictive unsupervised learning methods such as BYOL or SimSiam have shown impressive results, and counter-intuitively, do not collapse to trivial representations. In this work, we aim at exploring the simplest possible mathematical arguments towards explaining the underlying mechanisms behind self-predictive unsupervised learning. We start with the observation that those methods crucially rely on the presence of a predictor network (and stop-gradient). With simple linear algebra, we show that when using a linear predictor, the optimal predictor is close to an orthogonal projection, and propose a general framework based on orthonormalization that enables to interpret and give intuition on why BYOL works. In addition, this framework demonstrates the crucial role of the exponential moving average and stop-gradient operator in BYOL as an efficient orthonormalization mechanism. We use these insights to propose four new \emph{closed-form predictor} variants of BYOL to support our analysis. Our closed-form predictors outperform standard linear trainable predictor BYOL at $100$ and $300$ epochs (top-$1$ linear accuracy on ImageNet).
翻译:自预测无监督学习方法(如BYOL或SimSiam)展现出令人瞩目的性能,且反直觉地未坍缩至平凡表征。本研究旨在探索最简数学论证,以阐释自预测无监督学习背后的核心机制。我们首先观察到这类方法关键依赖于预测器网络(以及梯度停止)的存在。通过基础线性代数,我们证明当使用线性预测器时,其最优形式趋近于正交投影,并基于正交归一化提出通用框架,用以解释并直观理解BYOL的有效性。此外,该框架揭示了指数移动平均与梯度停止算子作为高效正交归一化机制的关键作用。基于这些洞见,我们提出四种BYOL的闭合形式预测器变体以支撑分析。在ImageNet上,我们的闭合形式预测器在100和300个训练周期中,其top-1线性准确率均超越标准线性可训练预测器BYOL。