Implicit regularization is an important way to interpret neural networks. Recent theory starts to explain implicit regularization with the model of deep matrix factorization (DMF) and analyze the trajectory of discrete gradient dynamics in the optimization process. These discrete gradient dynamics are relatively small but not infinitesimal, thus fitting well with the practical implementation of neural networks. Currently, discrete gradient dynamics analysis has been successfully applied to shallow networks but encounters the difficulty of complex computation for deep networks. In this work, we introduce another discrete gradient dynamics approach to explain implicit regularization, i.e. landscape analysis. It mainly focuses on gradient regions, such as saddle points and local minima. We theoretically establish the connection between saddle point escaping (SPE) stages and the matrix rank in DMF. We prove that, for a rank-R matrix reconstruction, DMF will converge to a second-order critical point after R stages of SPE. This conclusion is further experimentally verified on a low-rank matrix reconstruction problem. This work provides a new theory to analyze implicit regularization in deep learning.
翻译:隐式正则化是解释神经网络的重要方式。现有理论开始通过深度矩阵分解模型来阐述隐式正则化,并分析优化过程中离散梯度动力学的轨迹。这些离散梯度动力学虽较小但非无穷小,因此与神经网络的实践实现高度契合。目前,离散梯度动力学分析已成功应用于浅层网络,但针对深层网络面临计算复杂度高的困难。本研究提出另一种解释隐式正则化的离散梯度动力学方法,即景观分析。该方法主要关注梯度区域,如鞍点和局部最小值。我们从理论上建立了深度矩阵分解中逃逸鞍点阶段与矩阵秩之间的联系。证明对于秩为R的矩阵重构问题,经过R次逃逸鞍点阶段后,深度矩阵分解将收敛至二阶临界点。该结论在低秩矩阵重构问题上得到了实验验证。本研究为深度学习中的隐式正则化分析提供了新理论。