The study of Neural Tangent Kernels (NTKs) has provided much needed insight into convergence and generalization properties of neural networks in the over-parametrized (wide) limit by approximating the network using a first-order Taylor expansion with respect to its weights in the neighborhood of their initialization values. This allows neural network training to be analyzed from the perspective of reproducing kernel Hilbert spaces (RKHS), which is informative in the over-parametrized regime, but a poor approximation for narrower networks as the weights change more during training. Our goal is to extend beyond the limits of NTK toward a more general theory. We construct an exact power-series representation of the neural network in a finite neighborhood of the initial weights as an inner product of two feature maps, respectively from data and weight-step space, to feature space, allowing neural network training to be analyzed from the perspective of reproducing kernel {\em Banach} space (RKBS). We prove that, regardless of width, the training sequence produced by gradient descent can be exactly replicated by regularized sequential learning in RKBS. Using this, we present novel bound on uniform convergence where the iterations count and learning rate play a central role, giving new theoretical insight into neural network training.
翻译:神经正切核(NTK)的研究通过使用网络权重的初始化邻域内一阶泰勒展开来近似网络,为超参数化(宽)极限下神经网络的收敛性和泛化性质提供了重要见解。这允许从再生核希尔伯特空间(RKHS)的角度分析神经网络训练,该视角在超参数化条件下具有信息价值,但对于较窄的网络而言近似效果较差,因为训练过程中权重变化更大。我们的目标是超越NTK的局限性,建立更通用的理论。我们在初始权重的有限邻域内构建了神经网络的精确幂级数表示,该表示表示为两个特征映射的内积(分别来自数据空间和权重步长空间到特征空间),从而允许从再生核巴拿赫空间(RKBS)的角度分析神经网络训练。我们证明,无论网络宽度如何,梯度下降产生的训练序列都能被RKBS中的正则化序贯学习精确复制。基于此,我们提出了新的均匀收敛界,其中迭代次数和学习率扮演核心角色,为神经网络训练提供了新的理论洞见。