Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/\eta$ (where $\eta$ is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below $2/\eta$. While many other well-understood nonconvex objectives such as matrix factorization or two-layer networks can also converge despite large sharpness, there is often a larger gap between sharpness of the endpoint and $2/\eta$. In this paper, we study EoS phenomenon by constructing a simple function that has the same behavior. We give rigorous analysis for its training dynamics in a large local region and explain why the final converging point has sharpness close to $2/\eta$. Globally we observe that the training dynamics for our example has an interesting bifurcating behavior, which was also observed in the training of neural nets.
翻译:最近,研究人员观察到深度神经网络的梯度下降运行在“边缘稳定性”(EoS)机制中:锐度(Hessian矩阵的最大特征值)通常大于稳定性阈值$2/\eta$(其中$\eta$是步长)。尽管如此,损失函数仍会振荡并在长期内收敛,且最终锐度略低于$2/\eta$。尽管许多其他易于理解的非凸目标(如矩阵分解或两层网络)也能在锐度较大时收敛,但终点锐度与$2/\eta$之间往往存在较大差距。本文通过构造一个具有相同行为的简单函数来研究EoS现象。我们在大局部区域内对其训练动力学进行了严格分析,并解释了最终收敛点的锐度为何接近$2/\eta$。从全局来看,我们观察到示例的训练动力学具有有趣的分叉行为,这一现象在神经网络训练中也曾被观察到。