This paper studies the qualitative behavior and robustness of two variants of Minimal Random Code Learning (MIRACLE) used to compress variational Bayesian neural networks. MIRACLE implements a powerful, conditionally Gaussian variational approximation for the weight posterior $Q_{\mathbf{w}}$ and uses relative entropy coding to compress a weight sample from the posterior using a Gaussian coding distribution $P_{\mathbf{w}}$. To achieve the desired compression rate, $D_{\mathrm{KL}}[Q_{\mathbf{w}} \Vert P_{\mathbf{w}}]$ must be constrained, which requires a computationally expensive annealing procedure under the conventional mean-variance (Mean-Var) parameterization for $Q_{\mathbf{w}}$. Instead, we parameterize $Q_{\mathbf{w}}$ by its mean and KL divergence from $P_{\mathbf{w}}$ to constrain the compression cost to the desired value by construction. We demonstrate that variational training with Mean-KL parameterization converges twice as fast and maintains predictive performance after compression. Furthermore, we show that Mean-KL leads to more meaningful variational distributions with heavier tails and compressed weight samples which are more robust to pruning.
翻译:本文研究了用于压缩变分贝叶斯神经网络的最小随机编码学习(MIRACLE)两种变体的定性行为与鲁棒性。MIRACLE通过条件高斯变分近似实现权重重后验$Q_{\mathbf{w}}$的强效表达,并利用相对熵编码从后验中抽取权重样本,以高斯编码分布$P_{\mathbf{w}}$进行压缩。为达到目标压缩率,须约束$D_{\mathrm{KL}}[Q_{\mathbf{w}} \Vert P_{\mathbf{w}}]$,这在常规均值-方差(Mean-Var)参数化下需要代价高昂的退火过程。为此,我们提出通过均值和相对于$P_{\mathbf{w}}$的KL散度来参数化$Q_{\mathbf{w}}$,从而在构造上约束压缩成本至目标值。实验证明,采用均值-KL参数化的变分训练收敛速度提升两倍,且压缩后预测性能保持不变。此外,我们揭示均值-KL参数化能产生更具物理意义的变分分布——具有更重的尾部,且压缩后的权重样本对剪枝更具鲁棒性。