Figuring out small molecule binding sites in target proteins, in the resolution of either pocket or residue, is critical in many virtual and real drug-discovery scenarios. Since it is not always easy to find such binding sites based on domain knowledge or traditional methods, different deep learning methods that predict binding sites out of protein structures have been developed in recent years. Here we present a new such deep learning algorithm, that significantly outperformed all state-of-the-art baselines in terms of the both resolutions$\unicode{x2013}$pocket and residue. This good performance was also demonstrated in a case study involving the protein human serum albumin and its binding sites. Our algorithm included new ideas both in the model architecture and in the training method. For the model architecture, it incorporated SE(3)-invariant geometric self-attention layers that operate on top of residue-level CNN outputs. This residue-level processing of the model allowed a transfer learning between the two resolutions, which turned out to significantly improve the binding pocket prediction. Moreover, we developed novel augmentation method based on protein homology, which prevented our model from over-fitting. Overall, we believe that our contribution to the literature is twofold. First, we provided a new computational method for binding site prediction that is relevant to real-world applications, as shown by the good performance on different benchmarks and case study. Second, the novel ideas in our method$\unicode{x2013}$the model architecture, transfer learning and the homology augmentation$\unicode{x2013}$would serve as useful components in future works.
翻译:确定目标蛋白质中小分子结合位点(以口袋或残基为分辨率)是许多虚拟及真实药物发现场景中的关键环节。由于基于领域知识或传统方法寻找此类结合位点并非易事,近年来已发展出多种基于蛋白质结构预测结合位点的深度学习方法。本文提出一种新型深度学习算法,在口袋和残基两种分辨率下均显著超越所有当前最优基线模型。这一优异性能在涉及人血清白蛋白及其结合位点的案例研究中得到验证。我们的算法在模型架构和训练方法两方面均包含创新思路:模型架构方面,在残基层级CNN输出之上引入了SE(3)不变几何自注意力层;这种残基层级处理使两种分辨率间可实现迁移学习,显著提升了结合口袋预测性能。此外,我们基于蛋白质同源性开发了新型增强方法,有效防止了模型过拟合。总体而言,我们认为本研究对文献的贡献体现在两方面:首先,提供了与真实应用场景相关的结合位点预测新计算方法(不同基准测试和案例研究中的优异性能已证明其有效性);其次,方法中的创新要素——模型架构、迁移学习与同源增强——将为未来研究提供重要组件。