In this paper, we rigorously derive Central Limit Theorems (CLT) for Bayesian two-layerneural networks in the infinite-width limit and trained by variational inference on a regression task. The different networks are trained via different maximization schemes of the regularized evidence lower bound: (i) the idealized case with exact estimation of a multiple Gaussian integral from the reparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling, commonly known as Bayes-by-Backprop, and (iii) a computationally cheaper algorithm named Minimal VI. The latter was recently introduced by leveraging the information obtained at the level of the mean-field limit. Laws of large numbers are already rigorously proven for the three schemes that admits the same asymptotic limit. By deriving CLT, this work shows that the idealized and Bayes-by-Backprop schemes have similar fluctuation behavior, that is different from the Minimal VI one. Numerical experiments then illustrate that the Minimal VI scheme is still more efficient, in spite of bigger variances, thanks to its important gain in computational complexity.
翻译:本文针对无限宽度极限下、通过变分推断在回归任务上训练的贝叶斯双层神经网络,严格推导了中心极限定理(CLT)。不同网络通过正则化证据下界(ELBO)的三种最大化方案进行训练:(i)基于重参数化技巧精确估计多重高斯积分的理想化方案;(ii)采用蒙特卡洛采样的迷你批次方案(通常称为Bayes-by-Backprop);(iii)利用平均场极限层面信息构建的计算高效算法Minimal VI。针对这三种具有相同渐近极限的方案,大数定律已获严格证明。通过推导CLT,本研究表明理想化方案与Bayes-by-Backprop方案具有相似的波动特性,而Minimal VI方案则呈现不同特征。数值实验进一步表明,尽管Minimal VI方案具有更大的方差,但其在计算复杂度上的显著优势仍使其更具效率。