Recent advances in graph neural networks (GNNs) have allowed molecular simulations with accuracy on par with conventional gold-standard methods at a fraction of the computational cost. Nonetheless, as the field has been progressing to bigger and more complex architectures, state-of-the-art GNNs have become largely prohibitive for many large-scale applications. In this paper, we, for the first time, explore the utility of knowledge distillation (KD) for accelerating molecular GNNs. To this end, we devise KD strategies that facilitate the distillation of hidden representations in directional and equivariant GNNs and evaluate their performance on the regression task of energy and force prediction. We validate our protocols across different teacher-student configurations and demonstrate that they can boost the predictive accuracy of student models without altering their architecture. We also conduct comprehensive optimization of various components of our framework, and investigate the potential of data augmentation to further enhance performance. All in all, we manage to close as much as 59% of the gap in predictive accuracy between models like GemNet-OC and PaiNN with zero additional cost at inference.
翻译:近期图神经网络(GNN)的进展使得分子模拟在计算成本大幅降低的同时,能够达到与传统金标准方法相当的精度。然而,随着该领域向更大规模、更复杂的架构发展,最先进的GNN在许多大规模应用中已变得非常昂贵。本文首次探索了知识蒸馏(KD)在加速分子GNN中的应用价值。为此,我们设计了针对方向性和等变性GNN的隐藏表示蒸馏策略,并在能量和力预测的回归任务上评估其性能。我们在不同教师-学生配置下验证了所提出的方案,证明其能在不改变学生模型架构的前提下提升预测精度。我们还对框架中的各个组件进行了全面优化,并研究了数据增强进一步提升性能的潜力。总体而言,我们在推理零额外成本的情况下,成功将GemNet-OC与PaiNN等模型间的预测精度差距缩小了高达59%。