We study the stability and convergence of training deep ResNets with gradient descent. Specifically, we show that the parametric branch in the residual block should be scaled down by a factor $\tau =O(1/\sqrt{L})$ to guarantee stable forward/backward process, where $L$ is the number of residual blocks. Moreover, we establish a converse result that the forward process is unbounded when $\tau>L^{-\frac{1}{2}+c}$, for any positive constant $c$. The above two results together establish a sharp value of the scaling factor in determining the stability of deep ResNet. Based on the stability result, we further show that gradient descent finds the global minima if the ResNet is properly over-parameterized, which significantly improves over the previous work with a much larger range of $\tau$ that admits global convergence. Moreover, we show that the convergence rate is independent of the depth, theoretically justifying the advantage of ResNet over vanilla feedforward network. Empirically, with such a factor $\tau$, one can train deep ResNet without normalization layer. Moreover, for ResNets with normalization layer, adding such a factor $\tau$ also stabilizes the training and obtains significant performance gain for deep ResNet.
翻译:摘要:我们研究了使用梯度下降训练深度残差网络的稳定性与收敛性。具体而言,我们发现残差块中的参数分支应乘以缩放因子 $\tau =O(1/\sqrt{L})$ 以保证前向/反向过程的稳定性,其中 $L$ 是残差块的数量。此外,我们证明了对于任意正常数 $c$,当 $\tau>L^{-\frac{1}{2}+c}$ 时前向过程无界。上述两个结论共同确定了决定深度残差网络稳定性的缩放因子尖锐阈值。基于该稳定性结果,我们进一步证明了当残差网络适当过参数化时,梯度下降能够收敛到全局最小值,这显著改进了先前工作中允许全局收敛的 $\tau$ 取值区间。同时,我们证明了收敛速度与网络深度无关,从理论上验证了残差网络相较于普通前馈网络的优越性。实验表明,使用该缩放因子 $\tau$ 可在无归一化层的情况下训练深度残差网络。此外,对于带有归一化层的残差网络,添加该因子 $\tau$ 同样能稳定训练过程,并在深度网络中取得显著的性能提升。