In gradient descent dynamics of neural networks, the top eigenvalue of the Hessian of the loss (sharpness) displays a variety of robust phenomena throughout training. This includes early time regimes where the sharpness may decrease during early periods of training (sharpness reduction), and later time behavior such as progressive sharpening and edge of stability. We demonstrate that a simple $2$-layer linear network (UV model) trained on a single training example exhibits all of the essential sharpness phenomenology observed in real-world scenarios. By analyzing the structure of dynamical fixed points in function space and the vector field of function updates, we uncover the underlying mechanisms behind these sharpness trends. Our analysis reveals (i) the mechanism behind early sharpness reduction and progressive sharpening, (ii) the required conditions for edge of stability, and (iii) a period-doubling route to chaos on the edge of stability manifold as learning rate is increased. Finally, we demonstrate that various predictions from this simplified model generalize to real-world scenarios and discuss its limitations.
翻译:在神经网络梯度下降动力学中,损失函数海森矩阵的最大特征值(锐度)在整个训练过程中呈现出多种鲁棒现象。这包括早期训练阶段锐度可能下降的锐度缩减现象,以及渐进锐化与稳定性边缘等后期行为。我们证明,在单一训练样本上训练的简单双层线性网络(UV模型)能够展现现实场景中观察到的所有关键锐度现象。通过分析函数空间中动力学不动点的结构以及函数更新的向量场,我们揭示了这些锐度趋势背后的潜在机制。我们的分析阐明了:(i) 早期锐度缩减与渐进锐化的机制,(ii) 稳定性边缘成立的必要条件,以及 (iii) 随着学习率增加,稳定性边缘流形上通向混沌的倍周期分岔路径。最后,我们验证了该简化模型的多种预测对现实场景的推广性,并讨论了其局限性。