The Lipschitz constant is an important quantity that arises in analysing the convergence of gradient-based optimization methods. It is generally unclear how to estimate the Lipschitz constant of a complex model. Thus, this paper studies an important problem that may be useful to the broader area of non-convex optimization. The main result provides a local upper bound on the Lipschitz constants of a multi-layer feed-forward neural network and its gradient. Moreover, lower bounds are established as well, which are used to show that it is impossible to derive global upper bounds for the Lipschitz constants. In contrast to previous works, we compute the Lipschitz constants with respect to the network parameters and not with respect to the inputs. These constants are needed for the theoretical description of many step size schedulers of gradient based optimization schemes and their convergence analysis. The idea is both simple and effective. The results are extended to a generalization of neural networks, continuously deep neural networks, which are described by controlled ODEs.
翻译:利普希茨常数是分析基于梯度优化方法收敛性的重要量。目前尚不清楚如何估计复杂模型的利普希茨常数。因此,本文研究了一个可能对更广泛的非凸优化领域有所裨益的关键问题。主要结果给出了多层前馈神经网络及其梯度的利普希茨常数的局部上界。此外,本文还建立了下界,并用以证明无法推导出利普希茨常数的全局上界。与以往研究不同,我们计算的是关于网络参数的利普希茨常数,而非关于输入参数。这些常数对于许多基于梯度的优化方案中步长调度器的理论描述及其收敛性分析必不可少。该思想既简单又有效。研究结果进一步推广至神经网络的泛化形式——连续深度神经网络,此类网络由受控常微分方程描述。