We introduce Variational Joint Embedding (VJE), a reconstruction-free latent-variable framework for non-contrastive self-supervised learning in representation space. VJE maximizes a symmetric conditional evidence lower bound (ELBO) on paired encoder embeddings by defining a conditional likelihood directly on target representations, rather than optimizing a pointwise compatibility objective. The likelihood is instantiated as a heavy-tailed Student--\(t\) distribution on a polar representation of the target embedding, where a directional--radial decomposition separates angular agreement from magnitude consistency and mitigates norm-induced pathologies. The directional factor operates on the unit sphere, yielding a valid variational bound for the associated spherical subdensity model. An amortized inference network parameterizes a diagonal Gaussian posterior whose feature-wise variances are shared with the directional likelihood, yielding anisotropic uncertainty without auxiliary projection heads. Across ImageNet-1K, CIFAR-10/100, and STL-10, VJE is competitive with standard non-contrastive baselines under linear and \(k\)-NN evaluation, while providing probabilistic semantics directly in representation space for downstream uncertainty-aware applications. We validate these semantics through out-of-distribution detection, where representation-space likelihoods yield strong empirical performance. These results position the framework as a principled variational formulation of non-contrastive learning, in which structured feature-wise uncertainty is represented directly in the learned embedding space.
翻译:我们提出了变分联合嵌入(VJE),一种用于表示空间中非对比自监督学习的免重建潜变量框架。VJE通过直接在目标表示上定义条件似然(而非优化逐点兼容性目标),最大化配对编码器嵌入的对称条件证据下界(ELBO)。该似然基于目标嵌入的极坐标表示实现为重尾Student-t分布,其中方向-径向分解将角度一致性从幅度一致性中分离,并缓解了由范数引起的病态问题。方向因子在单位球面上运作,为关联的球面子密度模型提供了有效的变分界。一个摊销推理网络参数化对角高斯后验分布,其逐特征方差与方向似然共享,从而无需辅助投影头即可实现各向异性不确定性。在ImageNet-1K、CIFAR-10/100和STL-10数据集上,VJE在线性评估和k-NN评估中均达到与标准非对比基线相当的性能,同时直接在表示空间中为下游不确定性感知应用提供概率语义。我们通过分布外检测验证了这些语义,其中表示空间似然展现出强大的经验性能。这些结果将该框架定位为非对比学习的理论严谨变分公式,其中结构化的逐特征不确定性直接在学习的嵌入空间中得到表示。