Bayesian neural networks (BNNs) offer uncertainty quantification but come with the downside of substantially increased training and inference costs. Sparse BNNs have been investigated for efficient inference, typically by either slowly introducing sparsity throughout the training or by post-training compression of dense BNNs. The dilemma of how to cut down massive training costs remains, particularly given the requirement to learn about the uncertainty. To solve this challenge, we introduce Sparse Subspace Variational Inference (SSVI), the first fully sparse BNN framework that maintains a consistently highly sparse Bayesian model throughout the training and inference phases. Starting from a randomly initialized low-dimensional sparse subspace, our approach alternately optimizes the sparse subspace basis selection and its associated parameters. While basis selection is characterized as a non-differentiable problem, we approximate the optimal solution with a removal-and-addition strategy, guided by novel criteria based on weight distribution statistics. Our extensive experiments show that SSVI sets new benchmarks in crafting sparse BNNs, achieving, for instance, a 10-20x compression in model size with under 3\% performance drop, and up to 20x FLOPs reduction during training compared with dense VI training. Remarkably, SSVI also demonstrates enhanced robustness to hyperparameters, reducing the need for intricate tuning in VI and occasionally even surpassing VI-trained dense BNNs on both accuracy and uncertainty metrics.
翻译:贝叶斯神经网络(BNNs)提供了不确定性量化,但伴随着训练和推理成本显著增加的缺点。稀疏BNNs已被研究用于高效推理,通常通过在训练过程中缓慢引入稀疏性或对稠密BNNs进行训练后压缩。如何大幅降低训练成本这一难题依然存在,特别是在需要学习不确定性的背景下。为解决这一挑战,我们提出稀疏子空间变分推断(SSVI),这是首个在训练和推理阶段均保持高度稀疏贝叶斯模型的完全稀疏BNN框架。从随机初始化的低维稀疏子空间开始,我们的方法交替优化稀疏子空间基选择及其关联参数。由于基选择是一个不可微问题,我们通过基于权重分布统计的新型准则引导的移除-添加策略来近似最优解。大量实验表明,SSVI在构建稀疏BNNs方面树立了新的基准,例如在模型性能下降低于3%的情况下实现10-20倍的模型大小压缩,以及与稠密VI训练相比最高可达20倍的训练FLOPs缩减。值得注意的是,SSVI还展现出对超参数的更强鲁棒性,减少了VI中复杂调参的需求,有时甚至在准确性和不确定性指标上超越VI训练的稠密BNNs。