How to conduct teacher training for knowledge distillation is still an open problem. It has been widely observed that a best-performing teacher does not necessarily yield the best-performing student, suggesting a fundamental discrepancy between the current teacher training practice and the ideal teacher training strategy. To fill this gap, we explore the feasibility of training a teacher that is oriented toward student performance with empirical risk minimization (ERM). Our analyses are inspired by the recent findings that the effectiveness of knowledge distillation hinges on the teacher's capability to approximate the true label distribution of training inputs. We theoretically establish that the ERM minimizer can approximate the true label distribution of training data as long as the feature extractor of the learner network is Lipschitz continuous and is robust to feature transformations. In light of our theory, we propose a teacher training method SoTeacher which incorporates Lipschitz regularization and consistency regularization into ERM. Experiments on benchmark datasets using various knowledge distillation algorithms and teacher-student pairs confirm that SoTeacher can improve student accuracy consistently.
翻译:如何为知识蒸馏进行教师训练仍是一个待解决的问题。广泛观察到,性能最佳的教师并不一定能培养出性能最佳的学生,这表明当前教师训练实践与理想教师训练策略之间存在根本性差异。为填补这一空白,我们探索了通过经验风险最小化训练面向学生性能的教师的可行性。我们的分析受到最近发现的启发:知识蒸馏的有效性取决于教师逼近训练输入真实标签分布的能力。我们从理论上证明,只要学习网络的特征提取器是Lipschitz连续的且对特征变换具有鲁棒性,经验风险最小化最小化器就能逼近训练数据的真实标签分布。基于我们的理论,我们提出了一种教师训练方法SoTeacher,它将Lipschitz正则化和一致性正则化融入经验风险最小化中。使用多种知识蒸馏算法和师生对在基准数据集上的实验证实,SoTeacher能够持续提升学生的准确率。