Neural networks with wide layers have attracted significant attention due to their equivalence to Gaussian processes, enabling perfect fitting of training data while maintaining generalization performance, known as benign overfitting. However, existing results mainly focus on shallow or finite-depth networks, necessitating a comprehensive analysis of wide neural networks with infinite-depth layers, such as neural ordinary differential equations (ODEs) and deep equilibrium models (DEQs). In this paper, we specifically investigate the deep equilibrium model (DEQ), an infinite-depth neural network with shared weight matrices across layers. Our analysis reveals that as the width of DEQ layers approaches infinity, it converges to a Gaussian process, establishing what is known as the Neural Network and Gaussian Process (NNGP) correspondence. Remarkably, this convergence holds even when the limits of depth and width are interchanged, which is not observed in typical infinite-depth Multilayer Perceptron (MLP) networks. Furthermore, we demonstrate that the associated Gaussian vector remains non-degenerate for any pairwise distinct input data, ensuring a strictly positive smallest eigenvalue of the corresponding kernel matrix using the NNGP kernel. These findings serve as fundamental elements for studying the training and generalization of DEQs, laying the groundwork for future research in this area.
翻译:具有宽层的神经网络因其与高斯过程的等价性而受到广泛关注,这使得它们能够在保持泛化性能的同时完美拟合训练数据,这种现象被称为良性过拟合。然而,现有结果主要集中于浅层或有限深度网络,需要对具有无限深度层的宽神经网络(如神经常微分方程(ODE)和深度平衡模型(DEQ))进行综合分析。本文专门研究了深度平衡模型(DEQ),这是一种跨层共享权重矩阵的无限深度神经网络。我们的分析表明,当DEQ层的宽度趋近于无穷大时,它会收敛到高斯过程,从而建立了所谓的神经网络与高斯过程(NNGP)对应关系。值得注意的是,即使深度和宽度的极限顺序互换,这种收敛仍然成立,而这在典型的无限深度多层感知机(MLP)网络中并未观察到。此外,我们证明了对于任意两两不同的输入数据,相关的高斯向量保持非退化性,从而确保了使用NNGP核的对应核矩阵的最小特征值严格为正。这些发现为研究DEQ的训练与泛化奠定了基本要素,为该领域的未来研究奠定了基础。