Recent findings reveal that over-parameterized deep neural networks, trained beyond zero training-error, exhibit a distinctive structural pattern at the final layer, termed as Neural-collapse (NC). These results indicate that the final hidden-layer outputs in such networks display minimal within-class variations over the training set. While existing research extensively investigates this phenomenon under cross-entropy loss, there are fewer studies focusing on its contrastive counterpart, supervised contrastive (SC) loss. Through the lens of NC, this paper employs an analytical approach to study the solutions derived from optimizing the SC loss. We adopt the unconstrained features model (UFM) as a representative proxy for unveiling NC-related phenomena in sufficiently over-parameterized deep networks. We show that, despite the non-convexity of SC loss minimization, all local minima are global minima. Furthermore, the minimizer is unique (up to a rotation). We prove our results by formalizing a tight convex relaxation of the UFM. Finally, through this convex formulation, we delve deeper into characterizing the properties of global solutions under label-imbalanced training data.
翻译:近期研究发现,在达到零训练误差后继续训练的过参数化深度神经网络,其最终层会表现出一种独特的结构模式,即神经坍缩现象。这些结果表明,此类网络中最终隐藏层输出在训练集上的类内变异最小。尽管现有研究已广泛探讨交叉熵损失下的该现象,但针对其对比学习对应物——监督对比损失的同类研究仍较为有限。本文从神经坍缩视角出发,采用分析方法研究优化监督对比损失所得的解。我们采用无约束特征模型作为揭示充分过参数化深度网络中神经坍缩相关现象的代表性代理模型。研究表明,尽管监督对比损失最小化具有非凸性,但其所有局部极小值均为全局极小值,且极小值解(在旋转变换意义下)唯一。我们通过建立无约束特征模型的紧凑凸松弛形式证明了上述结论。最终,借助这一凸优化框架,我们深入刻画了标签非均衡训练数据下全局解的性质。