In this paper, we provide a rigorous proof of convergence of the Adaptive Moment Estimate (Adam) algorithm for a wide class of optimization objectives. Despite the popularity and efficiency of the Adam algorithm in training deep neural networks, its theoretical properties are not yet fully understood, and existing convergence proofs require unrealistically strong assumptions, such as globally bounded gradients, to show the convergence to stationary points. In this paper, we show that Adam provably converges to $\epsilon$-stationary points with ${O}(\epsilon^{-4})$ gradient complexity under far more realistic conditions. The key to our analysis is a new proof of boundedness of gradients along the optimization trajectory of Adam, under a generalized smoothness assumption according to which the local smoothness (i.e., Hessian norm when it exists) is bounded by a sub-quadratic function of the gradient norm. Moreover, we propose a variance-reduced version of Adam with an accelerated gradient complexity of ${O}(\epsilon^{-3})$.
翻译:本文针对一类广泛的优化目标,提供了自适应矩估计算法(Adam)收敛性的严格证明。尽管Adam算法在深度神经网络训练中广受欢迎且高效,但其理论性质尚未完全明晰。现有收敛性证明需依赖不切实际的强假设(如全局有界梯度)来证明其收敛至驻点。本文证明,在远更符合实际条件的假设下,Adam算法能以${O}(\epsilon^{-4})$的梯度复杂度确定性地收敛到$\epsilon$-驻点。我们分析的关键在于:在广义光滑性假设下(即局部光滑性(当Hessian矩阵存在时其范数)受梯度范数的次二次函数约束),提出了一种沿Adam优化轨迹的梯度有界性的新证明方法。此外,我们提出了Adam的方差缩减改进版本,其加速梯度复杂度可达${O}(\epsilon^{-3})$。