We discover restrained numerical instabilities in current training practices of deep networks with stochastic gradient descent (SGD). We show numerical error (on the order of the smallest floating point bit) induced from floating point arithmetic in training deep nets can be amplified significantly and result in significant test accuracy variance, comparable to the test accuracy variance due to stochasticity in SGD. We show how this is likely traced to instabilities of the optimization dynamics that are restrained, i.e., localized over iterations and regions of the weight tensor space. We do this by presenting a theoretical framework using numerical analysis of partial differential equations (PDE), and analyzing the gradient descent PDE of convolutional neural networks (CNNs). We show that it is stable only under certain conditions on the learning rate and weight decay. We show that rather than blowing up when the conditions are violated, the instability can be restrained. We show this is a consequence of the non-linear PDE associated with the gradient descent of the CNN, whose local linearization changes when over-driving the step size of the discretization, resulting in a stabilizing effect. We link restrained instabilities to the recently discovered Edge of Stability (EoS) phenomena, in which the stable step size predicted by classical theory is exceeded while continuing to optimize the loss and still converging. Because restrained instabilities occur at the EoS, our theory provides new predictions about the EoS, in particular, the role of regularization and the dependence on the network complexity.
翻译:我们发现当前使用随机梯度下降(SGD)训练深度网络的实践中存在受限制的数值不稳定性。研究表明,训练深度网络时浮点运算引起的数值误差(量级为最小浮点比特位)可能被显著放大,导致测试精度出现显著方差,其程度堪比SGD随机性造成的测试精度方差。我们证明这很可能源于优化动力学的受限制不稳定性,即这种不稳定性在迭代次数和权重张量空间区域上具有局部化特征。为此,我们提出一个基于偏微分方程(PDE)数值分析的理论框架,并分析了卷积神经网络(CNN)的梯度下降PDE。研究显示,该PDE仅在特定学习率和权重衰减条件下保持稳定。我们证明当条件被违反时,不稳定性并非直接爆发,而是呈现受限制特性。这源于CNN梯度下降对应的非线性PDE:当过度驱动离散化步长时,其局部线性化形式会发生改变,从而产生稳定化效应。我们将受限制不稳定性与近期发现的稳定边缘(EoS)现象联系起来——该现象中,经典理论预测的稳定步长虽被超越,但损失函数仍持续优化并最终收敛。由于受限制不稳定性发生于EoS区域,我们的理论为EoS提供了新预测,特别是正则化作用及其对网络复杂度的依赖关系。