While much progress has been achieved over the last decades in neuro-inspired machine learning, there are still fundamental theoretical problems in gradient-based learning using combinations of neurons. These problems, such as saddle points and suboptimal plateaus of the cost function, can lead in theory and practice to failures of learning. In addition, the discrete step size selection of the gradient is problematic since too large steps can lead to instability and too small steps slow down the learning. This paper describes an alternative discrete MinMax learning approach for continuous piece-wise linear functions. Global exponential convergence of the algorithm is established using Contraction Theory with Inequality Constraints, which is extended from the continuous to the discrete case in this paper: The parametrization of each linear function piece is, in contrast to deep learning, linear in the proposed MinMax network. This allows a linear regression stability proof as long as measurements do not transit from one linear region to its neighbouring linear region. The step size of the discrete gradient descent is Lagrangian limited orthogonal to the edge of two neighbouring linear functions. It will be shown that this Lagrangian step limitation does not decrease the convergence of the unconstrained system dynamics in contrast to a step size limitation in the direction of the gradient. We show that the convergence rate of a constrained piece-wise linear function learning is equivalent to the exponential convergence rates of the individual local linear regions.
翻译:尽管过去几十年在神经启发式机器学习领域取得了诸多进展,但基于梯度学习并结合神经元组合的方法仍存在根本性理论问题。这些问题(例如代价函数的鞍点和次优平台)在理论和实践中均可能导致学习失败。此外,梯度离散步长的选择也面临困境:步长过大会引发不稳定性,步长过小则会降低学习速度。本文提出了一种面向连续分段线性函数的替代性离散MinMax学习方法。通过引入带不等式约束的收缩理论(本文将其从连续情形扩展至离散情形),证明了该算法的全局指数收敛性:与深度网络不同,所提MinMax网络中每个线性函数分段的参数化具有线性特性。这保证了当测量值未从一个线性区域过渡至相邻线性区域时,可构建线性回归稳定性证明。离散梯度下降的步长受拉格朗日约束,其方向正交于两个相邻线性函数边缘。研究表明,与沿梯度方向的步长限制不同,这种拉格朗日步长限制不会降低无约束系统动力学的收敛性。我们证明,受约束的分段线性函数学习的收敛速率等价于各局部线性区域的指数收敛速率。