When training a neural network for classification, the feature vectors of the training set are known to collapse to the vertices of a regular simplex, provided the dimension $d$ of the feature space and the number $n$ of classes satisfies $n\leq d+1$. This phenomenon is known as neural collapse. For other applications like language models, one instead takes $n\gg d$. Here, the neural collapse phenomenon still occurs, but with different emergent geometric figures. We characterize these geometric figures in the orthoplex regime where $d+2\leq n\leq 2d$. The techniques in our analysis primarily involve Radon's theorem and convexity.
翻译:在训练用于分类任务的神经网络时,已知当特征空间维度 $d$ 与类别数量 $n$ 满足 $n\leq d+1$ 时,训练集的特征向量会坍缩至正则单纯形的顶点。这一现象被称为神经坍缩。对于语言模型等其他应用场景,通常需满足 $n\gg d$ 的条件。此时神经坍缩现象依然存在,但涌现出不同的几何构型。我们在 $d+2\leq n\leq 2d$ 的正交体状态下刻画了这些几何构型。分析过程中主要运用了Radon定理与凸性理论。