In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogeneous activations, even those with biases, residual connections, and normalization layers, while structured enough to enable geometric analysis of its gradient dynamics. Using this analysis, we generalize the existing results of maximum-margin bias for homogeneous networks to this richer class of models. We find that gradient flow implicitly favors a subset of the parameters, unlike in the case of a homogeneous model where all parameters are treated equally. We demonstrate through simple examples how this strong favoritism toward minimizing an asymmetric norm can degrade the robustness of quasi-homogeneous models. On the other hand, we conjecture that this norm-minimization discards, when possible, unnecessary higher-order parameters, reducing the model to a sparser parameterization. Lastly, by applying our theorem to sufficiently expressive neural networks with normalization layers, we reveal a universal mechanism behind the empirical phenomenon of Neural Collapse.
翻译:本研究探讨了在指数损失函数下,使用梯度流训练的准齐次神经网络在超过可分性点后的最大间隔偏差。我们提出了准齐次模型这一类别,该类模型具有足够的表达能力,能够描述几乎所有使用齐次激活函数的神经网络,包括那些包含偏置项、残差连接和归一化层的网络,同时其结构也足以对其梯度动力学进行几何分析。通过这一分析,我们将现有的齐次网络最大间隔偏差结果推广到这一更丰富的模型类别。我们发现,与齐次模型中所有参数被平等对待的情况不同,梯度流隐式地偏向于参数的某个子集。我们通过简单示例展示了这种强偏好于最小化非对称范数的行为如何会降低准齐次模型的鲁棒性。另一方面,我们推测这种范数最小化在可能的情况下会丢弃不必要的更高阶参数,从而将模型简化为更稀疏的参数化形式。最后,通过将我们的定理应用于具有归一化层的充分表达能力神经网络,我们揭示了神经坍缩这一经验现象背后的通用机制。