Deep learning approaches achieved significant progress in predicting protein structures. These methods are often applied to protein-protein interactions (PPIs) yet require Multiple Sequence Alignment (MSA) which is unavailable for various interactions, such as antibody-antigen. Computational docking methods are capable of sampling accurate complex models, but also produce thousands of invalid configurations. The design of scoring functions for identifying accurate models is a long-standing challenge. We develop a novel attention-based Graph Neural Network (GNN), ContactNet, for classifying PPI models obtained from docking algorithms into accurate and incorrect ones. When trained on docked antigen and modeled antibody structures, ContactNet doubles the accuracy of current state-of-the-art scoring functions, achieving accurate models among its Top-10 at 43% of the test cases. When applied to unbound antibodies, its Top-10 accuracy increases to 65%. This performance is achieved without MSA and the approach is applicable to other types of interactions, such as host-pathogens or general PPIs.
翻译:深度学习方法在预测蛋白质结构方面取得了显著进展。这些方法常被应用于蛋白质-蛋白质相互作用(PPI)预测,但通常需要多序列比对(MSA),而许多相互作用(如抗体-抗原相互作用)缺乏可用的MSA数据。计算对接方法能够采样生成精确的复合物模型,但同时也会产生数千个无效构型。设计用于识别精确模型的评分函数是一个长期存在的挑战。我们开发了一种新颖的基于注意力机制的图神经网络(GNN)——ContactNet,用于将对接算法获得的PPI模型分类为精确模型与错误模型。当使用对接的抗原与建模的抗体结构进行训练时,ContactNet将当前最先进评分函数的准确率提升了一倍,在43%的测试案例中,其Top-10预测结果包含精确模型。当应用于未结合抗体时,其Top-10准确率进一步提升至65%。这一性能的取得无需依赖MSA,且该方法可推广至其他类型的相互作用,如宿主-病原体相互作用或一般性PPI。