Neural Collapse refers to the curious phenomenon in the end of training of a neural network, where feature vectors and classification weights converge to a very simple geometrical arrangement (a simplex). While it has been observed empirically in various cases and has been theoretically motivated, its connection with crucial properties of neural networks, like their generalization and robustness, remains unclear. In this work, we study the stability properties of these simplices. We find that the simplex structure disappears under small adversarial attacks, and that perturbed examples "leap" between simplex vertices. We further analyze the geometry of networks that are optimized to be robust against adversarial perturbations of the input, and find that Neural Collapse is a pervasive phenomenon in these cases as well, with clean and perturbed representations forming aligned simplices, and giving rise to a robust simple nearest-neighbor classifier. By studying the propagation of the amount of collapse inside the network, we identify novel properties of both robust and non-robust machine learning models, and show that earlier, unlike later layers maintain reliable simplices on perturbed data.
翻译:神经坍缩指的是神经网络训练末期出现的一种奇特现象,即特征向量与分类权重收敛到一种极其简单的几何排列(单形结构)。尽管该现象在多种场景下被经验观测到并得到了理论解释,但其与神经网络关键特性(如泛化性与鲁棒性)的关联仍不明确。本文研究了这些单形结构的稳定性特性,发现该结构在小的对抗性攻击下会消失,且受扰动样本会在单形顶点间“跳跃”。我们进一步分析了针对输入扰动优化鲁棒性的网络几何结构,发现神经坍缩在这些情形下同样普遍存在——干净表征与扰动表征形成对齐的单形结构,并催生出鲁棒的简单最近邻分类器。通过研究坍缩程度在网络内部的传播规律,我们揭示了鲁棒与非鲁棒机器学习模型的新特性,并证明浅层(不同于深层)能够在受扰动数据上维持可靠的单形结构。