In knowledge distillation research, feature-based methods have dominated due to their ability to effectively tap into extensive teacher models. In contrast, logit-based approaches are considered to be less adept at extracting hidden 'dark knowledge' from teachers. To bridge this gap, we present LumiNet, a novel knowledge-transfer algorithm designed to enhance logit-based distillation. We introduce a perception matrix that aims to recalibrate logits through adjustments based on the model's representation capability. By meticulously analyzing intra-class dynamics, LumiNet reconstructs more granular inter-class relationships, enabling the student model to learn a richer breadth of knowledge. Both teacher and student models are mapped onto this refined matrix, with the student's goal being to minimize representational discrepancies. Rigorous testing on benchmark datasets (CIFAR-100, ImageNet, and MSCOCO) attests to LumiNet's efficacy, revealing its competitive edge over leading feature-based methods. Moreover, in exploring the realm of transfer learning, we assess how effectively the student model, trained using our method, adapts to downstream tasks. Notably, when applied to Tiny ImageNet, the transferred features exhibit remarkable performance, further underscoring LumiNet's versatility and robustness in diverse settings. With LumiNet, we hope to steer the research discourse towards a renewed interest in the latent capabilities of logit-based knowledge distillation.
翻译:在知识蒸馏研究中,基于特征的方法因其能够有效利用大型教师模型而占据主导地位。相比之下,基于逻辑的方法被认为在提取教师模型中隐藏的"暗知识"方面能力较弱。为弥补这一差距,我们提出LumiNet,一种旨在增强基于逻辑蒸馏的新型知识迁移算法。我们引入感知矩阵,通过基于模型表示能力的调整来重新校准逻辑值。通过细致分析类内动态,LumiNet重建了更精细的类间关系,使student模型能够学习更丰富的知识广度。教师模型和学生模型均被映射到此精细化矩阵上,学生模型的目标是最小化表示差异。在基准数据集(CIFAR-100、ImageNet和MSCOCO)上的严格测试证实了LumiNet的有效性,揭示了其相较于领先的基于特征方法的竞争优势。此外,在探索迁移学习领域时,我们评估了使用我们的方法训练的学生模型适应下游任务的效果。值得注意的是,当应用于Tiny ImageNet时,迁移后的特征展现出卓越性能,进一步凸显了LumiNet在不同场景下的多功能性和鲁棒性。通过LumiNet,我们希望引导研究界重新关注基于逻辑的知识蒸馏的潜在能力。