We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant $n$. Under mild assumptions on the non-linearity, we obtain quantitative bounds on normal approximations valid at large but finite $n$ and any fixed network depth. Our theorems show both for the finite-dimensional distributions and the entire process, that the distance between a random fully connected network (and its derivatives) to the corresponding infinite width Gaussian process scales like $n^{-\gamma}$ for $\gamma>0$, with the exponent depending on the metric used to measure discrepancy. Our bounds are strictly stronger in terms of their dependence on network width than any previously available in the literature; in the one-dimensional case, we also prove that they are optimal, i.e., we establish matching lower bounds.
翻译:我们研究具有随机高斯权重和偏置的全连接神经网络分布,其中隐藏层宽度与一个较大的常数$n$成比例。在对非线性函数施加温和假设的条件下,我们获得了适用于大而有限$n$及任意固定网络深度的正态逼近定量界。我们的定理表明,无论是对于有限维分布还是整个过程,随机全连接网络(及其导数)与对应无限宽度高斯过程之间的距离均按$n^{-\gamma}$($\gamma>0$)的比例缩放,其中指数取决于度量差异所采用的度量方式。在依赖网络宽度的维度上,我们给出的界严格强于文献中已有的任何结果;在一维情形中,我们还证明了这些界的最优性,即建立了匹配的下界。