Recent mainstream masked distillation methods function by reconstructing selectively masked areas of a student network from the feature map of its teacher counterpart. In these methods, the masked regions need to be properly selected, such that reconstructed features encode sufficient discrimination and representation capability like the teacher feature. However, previous masked distillation methods only focus on spatial masking, making the resulting masked areas biased towards spatial importance without encoding informative channel clues. In this study, we devise a Dual Masked Knowledge Distillation (DMKD) framework which can capture both spatially important and channel-wise informative clues for comprehensive masked feature reconstruction. More specifically, we employ dual attention mechanism for guiding the respective masking branches, leading to reconstructed feature encoding dual significance. Furthermore, fusing the reconstructed features is achieved by self-adjustable weighting strategy for effective feature distillation. Our experiments on object detection task demonstrate that the student networks achieve performance gains of 4.1% and 4.3% with the help of our method when RetinaNet and Cascade Mask R-CNN are respectively used as the teacher networks, while outperforming the other state-of-the-art distillation methods.
翻译:近期主流的掩码蒸馏方法通过从教师网络的特征图中重建学生网络选择性掩码区域来实现。在这些方法中,需合理选择掩码区域,使得重建后的特征能像教师特征一样编码足够的判别能力和表示能力。然而,以往的掩码蒸馏方法仅关注空间掩码,导致掩码区域偏向空间重要性,而未能编码信息丰富的通道线索。本研究提出了一种双重掩码知识蒸馏(DMKD)框架,该框架能够同时捕获空间重要性和通道信息丰富的线索,以实现全面的掩码特征重建。具体而言,我们采用双重注意力机制来引导各自的掩码分支,从而生成编码双重重要性的重建特征。此外,通过自调节权重策略实现重建特征的融合,以进行有效的特征蒸馏。我们在目标检测任务上的实验表明,当使用RetinaNet和Cascade Mask R-CNN分别作为教师网络时,学生网络借助我们的方法分别获得了4.1%和4.3%的性能提升,同时超越了其他最先进的蒸馏方法。