Network pruning techniques, including weight pruning and filter pruning, reveal that most state-of-the-art neural networks can be accelerated without a significant performance drop. This work focuses on filter pruning which enables accelerated inference with any off-the-shelf deep learning library and hardware. We propose the concept of \emph{network pruning spaces} that parametrize populations of subnetwork architectures. Based on this concept, we explore the structure aspect of subnetworks that result in minimal loss of accuracy in different pruning regimes and arrive at a series of observations by comparing subnetwork distributions. We conjecture through empirical studies that there exists an optimal FLOPs-to-parameter-bucket ratio related to the design of original network in a pruning regime. Statistically, the structure of a winning subnetwork guarantees an approximately optimal ratio in this regime. Upon our conjectures, we further refine the initial pruning space to reduce the cost of searching a good subnetwork architecture. Our experimental results on ImageNet show that the subnetwork we found is superior to those from the state-of-the-art pruning methods under comparable FLOPs.
翻译:网络剪枝技术(包括权重剪枝和滤波器剪枝)表明,大多数最先进的神经网络可以在不显著降低性能的情况下实现加速。本工作聚焦于滤波器剪枝,该方法能利用任何现成的深度学习库和硬件实现推理加速。我们提出"网络剪枝空间"的概念,该概念对子网络架构的种群进行参数化。基于这一概念,我们探索了在不同剪枝机制下导致精度损失最小的子网络结构特征,并通过比较子网络分布得出一系列观察结论。通过实证研究,我们推测在某一剪枝机制下,存在与原始网络设计相关的浮点运算次数与参数桶的优化比例。统计显示,优胜子网络的结构能够保证在该机制下近似达到此优化比例。基于该推论,我们进一步优化初始剪枝空间以降低搜索优质子网络架构的成本。我们在ImageNet上的实验结果表明,在相当的FLOPs条件下,我们所发现的子网络性能优于现有最先进剪枝方法得到的子网络。