We analyze Elman-type Recurrent Reural Networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large width limit. We also show that the fixed points of the limiting infinite-width dynamics are globally optimal, under some assumptions on the initialization of the weights. Our results establish optimality for feature-learning with wide RNNs in the mean-field regime
翻译:我们分析了Elman型递归神经网络(RNN)及其在平均场框架下的训练过程。具体而言,我们证明了在大宽度极限下,RNN的梯度下降训练动力学收敛到对应的平均场表示。我们还证明了,在权重初始化满足某些假设的条件下,极限无限宽动力学的固定点是全局最优的。我们的结果确立了在平均场框架下使用宽RNN进行特征学习的最优性。