Pruning before training enables the deployment of neural networks on smart devices. By retaining weights conducive to generalization, pruned networks can be accommodated on resource-constrained smart devices. It is commonly held that the distance on weight norms between the initialized and the fully-trained networks correlates with generalization performance. However, as we have uncovered, inconsistency between this metric and generalization during training processes, which poses an obstacle to determine the pruned structures on smart devices in advance. In this paper, we introduce the concept of the learning gap, emphasizing its accurate correlation with generalization. Experiments show that the learning gap, in the form of feature maps from the penultimate layer of networks, aligns with variations of generalization performance. We propose a novel learning framework, LNPT, which enables mature networks on the cloud to provide online guidance for network pruning and learning on smart devices with unlabeled data. Our results demonstrate the superiority of this approach over supervised training.
翻译:训练前进行剪枝能够支持神经网络在智能设备上的部署。通过保留有助于泛化的权重,剪枝后的网络可适配资源受限的智能设备。现有观点认为,初始化网络与完全训练网络之间的权重大小距离与泛化性能相关。然而,我们发现该指标在训练过程中与泛化性能存在不一致性,这对预先确定智能设备上的剪枝结构构成了障碍。本文引入了学习差距的概念,重点阐明其与泛化性能的准确关联性。实验表明,以网络倒数第二层特征图形式呈现的学习差距与泛化性能的变化相一致。我们提出了一种新颖的学习框架LNPT,该框架能使云端成熟网络为智能设备上的网络剪枝与学习提供在线指导(使用无标签数据)。实验结果证明了该方法相较于监督训练的优越性。