Prediction of protein-ligand interactions (PLI) plays a crucial role in drug discovery as it guides the identification and optimization of molecules that effectively bind to target proteins. Despite remarkable advances in deep learning-based PLI prediction, the development of a versatile model capable of accurately scoring binding affinity and conducting efficient virtual screening remains a challenge. The main obstacle in achieving this lies in the scarcity of experimental structure-affinity data, which limits the generalization ability of existing models. Here, we propose a viable solution to address this challenge by introducing a novel data augmentation strategy combined with a physics-informed graph neural network. The model showed significant improvements in both scoring and screening, outperforming task-specific deep learning models in various tests including derivative benchmarks, and notably achieving results comparable to the state-of-the-art performance based on distance likelihood learning. This demonstrates the potential of this approach to drug discovery.
翻译:蛋白质-配体相互作用(PLI)的预测在药物发现中至关重要,因为它能够指导有效结合靶标蛋白的分子的识别与优化。尽管基于深度学习的PLI预测取得了显著进展,但开发一种既能精准评分结合亲和力又能高效进行虚拟筛选的通用模型仍面临挑战。实现这一目标的主要障碍在于实验结构-亲和力数据的稀缺性,这限制了现有模型的泛化能力。本文提出了一种可行的解决方案,通过引入新型数据增强策略并结合物理信息图神经网络来应对这一挑战。该模型在评分和筛选两方面均展现出显著提升,在包括衍生基准测试在内的多项测试中优于任务特定的深度学习模型,并取得了与基于距离似然学习的当前最优方法相媲美的结果。这充分证明了该方法在药物发现领域的应用潜力。