Knowledge distillation (KD) has received much attention due to its success in compressing networks to allow for their deployment in resource-constrained systems. While the problem of adversarial robustness has been studied before in the KD setting, previous works overlook what we term the relative calibration of the student network with respect to its teacher in terms of soft confidences. In particular, we focus on two crucial questions with regard to a teacher-student pair: (i) do the teacher and student disagree at points close to correctly classified dataset examples, and (ii) is the distilled student as confident as the teacher around dataset examples? These are critical questions when considering the deployment of a smaller student network trained from a robust teacher within a safety-critical setting. To address these questions, we introduce a faithful imitation framework to discuss the relative calibration of confidences, as well as provide empirical and certified methods to evaluate the relative calibration of a student w.r.t. its teacher. Further, to verifiably align the relative calibration incentives of the student to those of its teacher, we introduce faithful distillation. Our experiments on the MNIST and Fashion-MNIST datasets demonstrate the need for such an analysis and the advantages of the increased verifiability of faithful distillation over alternative adversarial distillation methods.
翻译:知识蒸馏(KD)因其在压缩网络以部署于资源受限系统方面的成功而备受关注。尽管对抗鲁棒性在KD设置中已有研究,但先前工作忽视了学生网络相对于教师网络在软置信度方面的所谓相对校准问题。具体而言,我们聚焦于师生对的两个关键问题:(i)教师和学生是否在接近正确分类数据样本的点上存在分歧;(ii)在学生处理数据样本时,其置信度是否与教师相当?在考虑将基于鲁棒教师训练的更小型学生网络部署于安全关键场景时,这些问题至关重要。为解决这些问题,我们引入了一个忠实模仿框架来讨论置信度的相对校准,并提供经验性与可验证方法来评估学生相对于教师的相对校准性能。此外,为可验证地使学生与教师在相对校准激励上保持一致,我们提出了忠实蒸馏。在MNIST和Fashion-MNIST数据集上的实验表明,此类分析的必要性以及忠实蒸馏相较于其他对抗蒸馏方法在可验证性提升方面的优势。