Pre-trained models have been successful in many protein engineering tasks. Most notably, sequence-based models have achieved state-of-the-art performance on protein fitness prediction while structure-based models have been used experimentally to develop proteins with enhanced functions. However, there is a research gap in comparing structure- and sequence-based methods for predicting protein variants that are better than the wildtype protein. This paper aims to address this gap by conducting a comparative study between the abilities of equivariant graph neural networks (EGNNs) and sequence-based approaches to identify promising amino-acid mutations. The results show that our proposed structural approach achieves a competitive performance to sequence-based methods while being trained on significantly fewer molecules. Additionally, we find that combining assay labelled data with structure pre-trained models yields similar trends as with sequence pre-trained models. Our code and trained models can be found at: https://github.com/semiluna/partIII-amino-acid-prediction.
翻译:预训练模型已在许多蛋白质工程任务中取得成功。最值得注意的是,基于序列的模型在蛋白质适应度预测方面取得了最先进的性能,而基于结构的模型已被实验用于开发具有增强功能的蛋白质。然而,在比较基于结构和基于序列的方法以预测优于野生型蛋白质的蛋白质变体方面存在研究空白。本文旨在通过比较等变图神经网络(EGNNs)与基于序列的方法在识别有前途的氨基酸突变方面的能力来填补这一空白。结果表明,我们提出的结构方法在训练数据显著减少的情况下,达到了与基于序列方法相竞争的性能。此外,我们发现将检测标记数据与结构预训练模型相结合会产生与序列预训练模型相似的趋势。我们的代码和训练模型可在以下网址获取:https://github.com/semiluna/partIII-amino-acid-prediction。