From the classical and influential works of Neal (1996), it is known that the infinite width scaling limit of a Bayesian neural network with one hidden layer is a Gaussian process, \emph{when the network weights have bounded prior variance}. Neal's result has been extended to networks with multiple hidden layers and to convolutional neural networks, also with Gaussian process scaling limits. The tractable properties of Gaussian processes then allow straightforward posterior inference and uncertainty quantification, considerably simplifying the study of the limit process compared to a network of finite width. Neural network weights with unbounded variance, however, pose unique challenges. In this case, the classical central limit theorem breaks down and it is well known that the scaling limit is an $\alpha$-stable process under suitable conditions. However, current literature is primarily limited to forward simulations under these processes and the problem of posterior inference under such a scaling limit remains largely unaddressed, unlike in the Gaussian process case. To this end, our contribution is an interpretable and computationally efficient procedure for posterior inference, using a \emph{conditionally Gaussian} representation, that then allows full use of the Gaussian process machinery for tractable posterior inference and uncertainty quantification in the non-Gaussian regime.
翻译:基于Neal(1996)经典且具有影响力的工作,我们已知:当网络权重具有有界先验方差时,单隐藏层贝叶斯神经网络的无限宽极限服从高斯过程。这一结果已拓展至多隐藏层网络与卷积神经网络,同样获得高斯过程尺度极限。高斯过程的易处理特性使后验推断与不确定性量化变得简便,相较于有限宽网络,极限过程的研究由此显著简化。然而,具有无界方差的神经网络权重带来了独特挑战。此时经典中心极限定理失效,学界公认在适当条件下其尺度极限为α-稳定过程。但现有文献主要局限于这些过程的正向模拟,与高斯过程情形不同,其尺度极限下的后验推断问题尚未得到充分解决。为此,我们提出一种可解释且计算高效的后验推断方法,通过“条件高斯”表示,使得在非高斯场景下仍可充分利用高斯过程框架实现易处理的后验推断与不确定性量化。