Pruning before training enables the deployment of neural networks on smart devices. By retaining weights conducive to generalization, pruned networks can be accommodated on resource-constrained smart devices. It is commonly held that the distance on weight norms between the initialized and the fully-trained networks correlates with generalization performance. However, as we have uncovered, inconsistency between this metric and generalization during training processes, which poses an obstacle to determine the pruned structures on smart devices in advance. In this paper, we introduce the concept of the learning gap, emphasizing its accurate correlation with generalization. Experiments show that the learning gap, in the form of feature maps from the penultimate layer of networks, aligns with variations of generalization performance. We propose a novel learning framework, LNPT, which enables mature networks on the cloud to provide online guidance for network pruning and learning on smart devices with unlabeled data. Our results demonstrate the superiority of this approach over supervised training.
翻译:摘要:训练前进行剪枝使得神经网络能够在智能设备上部署。通过保留有助于泛化的权重,剪枝后的网络可以适配资源受限的智能设备。通常认为,初始化网络与完全训练网络之间的权重范数距离与泛化性能相关。然而,正如我们所发现的,这一指标在训练过程中与泛化之间存在不一致性,这给提前确定智能设备上的剪枝结构带来了障碍。本文引入了学习差距的概念,强调其与泛化之间的准确相关性。实验表明,以网络倒数第二层特征图形式呈现的学习差距与泛化性能的变化保持一致。我们提出了一种新型学习框架LNPT,该框架使得云端成熟网络能够为智能设备上的网络剪枝和学习提供在线指导,且无需使用标注数据。我们的结果表明,该方法优于监督式训练。