Recent mainstream masked distillation methods function by reconstructing selectively masked areas of a student network from the feature map of its teacher counterpart. In these methods, the masked regions need to be properly selected, such that reconstructed features encode sufficient discrimination and representation capability like the teacher feature. However, previous masked distillation methods only focus on spatial masking, making the resulting masked areas biased towards spatial importance without encoding informative channel clues. In this study, we devise a Dual Masked Knowledge Distillation (DMKD) framework which can capture both spatially important and channel-wise informative clues for comprehensive masked feature reconstruction. More specifically, we employ dual attention mechanism for guiding the respective masking branches, leading to reconstructed feature encoding dual significance. Furthermore, fusing the reconstructed features is achieved by self-adjustable weighting strategy for effective feature distillation. Our experiments on object detection task demonstrate that the student networks achieve performance gains of 4.1% and 4.3% with the help of our method when RetinaNet and Cascade Mask R-CNN are respectively used as the teacher networks, while outperforming the other state-of-the-art distillation methods.
翻译:近期主流掩码蒸馏方法通过从教师网络的特征图中选择性重建学生网络的掩码区域来发挥作用。在此类方法中,需合理选择掩码区域,使重建特征具备与教师特征相当的判别与表征能力。然而,现有掩码蒸馏方法仅关注空间维度的掩码操作,导致所得掩码区域偏向空间重要性而未能编码信息丰富的通道线索。本研究提出一种双掩码知识蒸馏框架,该框架能够同时捕捉空间重要线索与通道信息线索,实现全面的掩码特征重建。具体而言,我们采用双重注意力机制分别指导各掩码分支,使重建特征编码双重显著性。此外,通过自适应权重策略实现重建特征的融合,从而进行有效的特征蒸馏。在目标检测任务上的实验表明:当分别以RetinaNet和Cascade Mask R-CNN作为教师网络时,基于本方法的学生网络性能分别提升4.1%和4.3%,同时优于其他现有先进蒸馏方法。