Neural network compression has been an increasingly important subject, not only due to its practical relevance, but also due to its theoretical implications, as there is an explicit connection between compressibility and generalization error. Recent studies have shown that the choice of the hyperparameters of stochastic gradient descent (SGD) can have an effect on the compressibility of the learned parameter vector. These results, however, rely on unverifiable assumptions and the resulting theory does not provide a practical guideline due to its implicitness. In this study, we propose a simple modification for SGD, such that the outputs of the algorithm will be provably compressible without making any nontrivial assumptions. We consider a one-hidden-layer neural network trained with SGD, and show that if we inject additive heavy-tailed noise to the iterates at each iteration, for any compression rate, there exists a level of overparametrization such that the output of the algorithm will be compressible with high probability. To achieve this result, we make two main technical contributions: (i) we prove a 'propagation of chaos' result for a class of heavy-tailed stochastic differential equations, and (ii) we derive error estimates for their Euler discretization. Our experiments suggest that the proposed approach not only achieves increased compressibility with various models and datasets, but also leads to robust test performance under pruning, even in more realistic architectures that lie beyond our theoretical setting.
翻译:神经网络压缩已成为一个日益重要的课题,这不仅源于其实用价值,还在于其理论意义——可压缩性与泛化误差之间存在明确关联。最新研究表明,随机梯度下降(SGD)超参数的选择会影响学习参数向量的可压缩性。然而,这些结论依赖于无法验证的假设,且相关理论因缺乏显式指导而难以应用于实践。本研究提出一种针对SGD的简单改进方案,使得该算法输出在无需任何非平凡假设的前提下获得可证明的可压缩性。我们以单隐层神经网络在SGD训练下的表现为研究对象,证明若在每次迭代中向迭代过程注入加性重尾噪声,则对于任意压缩率,存在某个过参数化水平使得算法输出能以高概率实现可压缩。为获得这一结论,我们取得两项主要技术突破:(i)针对一类重尾随机微分方程证明了"混沌传播"性质,(ii)推导出该类方程欧拉离散化方法的误差估计。实验表明,该方案不仅能在多种模型与数据集上增强可压缩性,即便在超出理论框架的复杂架构中,剪枝后的测试性能依然保持稳健。