Stochastic-gradient sampling methods are often used to perform Bayesian inference on neural networks. It has been observed that the methods in which notions of differential geometry are included tend to have better performances, with the Riemannian metric improving posterior exploration by accounting for the local curvature. However, the existing methods often resort to simple diagonal metrics to remain computationally efficient. This loses some of the gains. We propose two non-diagonal metrics that can be used in stochastic-gradient samplers to improve convergence and exploration but have only a minor computational overhead over diagonal metrics. We show that for fully connected neural networks (NNs) with sparsity-inducing priors and convolutional NNs with correlated priors, using these metrics can provide improvements. For some other choices the posterior is sufficiently easy also for the simpler metrics.
翻译:随机梯度采样方法常被用于对神经网络进行贝叶斯推断。实验表明,融入微分几何概念的采样方法往往表现更优,黎曼度量通过考虑局部曲率来改进后验探索。然而,现有方法通常采用简单的对角度量以保证计算效率,这损失了部分优势。本文提出两种非对角度量,可在随机梯度采样器中用于提升收敛性和探索能力,且相较于对角度量仅增加极小的计算开销。我们证明:对于具有稀疏诱导先验的全连接神经网络和具有相关先验的卷积神经网络,使用这些度量能带来改进。而对于其他某些选择,即使是更简单的度量也足以充分刻画后验分布。