Network compression is now a mature sub-field of neural network research: over the last decade, significant progress has been made towards reducing the size of models and speeding up inference, while maintaining the classification accuracy. However, many works have observed that focusing on just the overall accuracy can be misguided. E.g., it has been shown that mismatches between the full and compressed models can be biased towards under-represented classes. This raises the important research question, can we achieve network compression while maintaining "semantic equivalence" with the original network? In this work, we study this question in the context of the "long tail" phenomenon in computer vision datasets observed by Feldman, et al. They argue that memorization of certain inputs (appropriately defined) is essential to achieving good generalization. As compression limits the capacity of a network (and hence also its ability to memorize), we study the question: are mismatches between the full and compressed models correlated with the memorized training data? We present positive evidence in this direction for image classification tasks, by considering different base architectures and compression schemes.
翻译:网络压缩如今已是神经网络研究中一个成熟的子领域:在过去十年中,在降低模型规模、加速推理的同时保持分类准确率方面取得了显著进展。然而,许多工作已指出,仅关注整体准确率可能具有误导性。例如,已有研究表明,完整模型与压缩模型之间的差异可能偏向于代表性不足的类别。这引发了一个重要的研究问题:我们能否在保持与原始网络“语义等价”的同时实现网络压缩?在本工作中,我们结合Feldman等人观察到的计算机视觉数据集中的“长尾”现象来研究这一问题。他们认为,对某些输入的适当定义的记忆对于实现良好的泛化能力至关重要。由于压缩限制了网络的容量(从而也限制了其记忆能力),我们研究以下问题:完整模型与压缩模型之间的差异是否与记忆的训练数据相关?通过考虑不同的基础架构和压缩方案,我们针对图像分类任务提供了支持这一方向的正面证据。