Stochastic-gradient sampling methods are often used to perform Bayesian inference on neural networks. It has been observed that the methods in which notions of differential geometry are included tend to have better performances, with the Riemannian metric improving posterior exploration by accounting for the local curvature. However, the existing methods often resort to simple diagonal metrics to remain computationally efficient. This loses some of the gains. We propose two non-diagonal metrics that can be used in stochastic-gradient samplers to improve convergence and exploration but have only a minor computational overhead over diagonal metrics. We show that for fully connected neural networks (NNs) with sparsity-inducing priors and convolutional NNs with correlated priors, using these metrics can provide improvements. For some other choices the posterior is sufficiently easy also for the simpler metrics.
翻译:随机梯度采样方法常用于对神经网络进行贝叶斯推断。研究表明,纳入微分几何概念的方法通常表现更优,其中黎曼度量通过考虑局部曲率来改进后验探索。然而,现有方法为保持计算效率常采用简单对角度量,这损失了部分增益。我们提出两种可用于随机梯度采样器的非对角度量,它们能在仅带来轻微计算开销的前提下改善收敛性与探索能力。实验证明,对于采用稀疏诱导先验的全连接神经网络和采用相关先验的卷积神经网络,使用这些度量可带来性能提升。而对于某些其他选择,后验分布足够简单,使得较简单的度量也能取得良好效果。