Knowledge distillation (KD) has shown great potential for transferring knowledge from a complex teacher model to a simple student model in which the heavy learning task can be accomplished efficiently and without losing too much prediction accuracy. Recently, many attempts have been made by applying the KD mechanism to the graph representation learning models such as graph neural networks (GNNs) to accelerate the model's inference speed via student models. However, many existing KD-based GNNs utilize MLP as a universal approximator in the student model to imitate the teacher model's process without considering the graph knowledge from the teacher model. In this work, we provide a KD-based framework on multi-scaled GNNs, known as graph framelet, and prove that by adequately utilizing the graph knowledge in a multi-scaled manner provided by graph framelet decomposition, the student model is capable of adapting both homophilic and heterophilic graphs and has the potential of alleviating the over-squashing issue with a simple yet effectively graph surgery. Furthermore, we show how the graph knowledge supplied by the teacher is learned and digested by the student model via both algebra and geometry. Comprehensive experiments show that our proposed model can generate learning accuracy identical to or even surpass the teacher model while maintaining the high speed of inference.
翻译:知识蒸馏(KD)在将复杂教师模型的知识迁移至简单学生模型方面展现出巨大潜力,使得学生模型能够高效完成繁重学习任务且不显著损失预测精度。近年来,研究者尝试将KD机制应用于图表示学习模型(如图神经网络GNN),通过学生模型加速模型推理速度。然而,现有基于KD的GNN方法多采用多层感知机(MLP)作为学生模型的通用逼近器来模仿教师模型的处理过程,却未充分考虑教师模型中的图结构知识。本研究提出一种基于KD的多尺度GNN框架——图框架(graph framelet),通过证明合理利用图框架分解提供的多尺度图知识,学生模型能够自适应同质性与异质性图,并具备通过简单而有效的图手术缓解过挤压问题的潜力。进一步,我们从代数和几何两个维度揭示了教师模型提供的图知识如何被学生模型学习与消化。综合实验表明,所提模型在保持高推理速度的同时,能够达到甚至超越教师模型的学习精度。