Directed evolution plays an indispensable role in protein engineering that revises existing protein sequences to attain new or enhanced functions. Accurately predicting the effects of protein variants necessitates an in-depth understanding of protein structure and function. Although large self-supervised language models have demonstrated remarkable performance in zero-shot inference using only protein sequences, these models inherently do not interpret the spatial characteristics of protein structures, which are crucial for comprehending protein folding stability and internal molecular interactions. This paper introduces a novel pre-training framework that cascades sequential and geometric analyzers for protein primary and tertiary structures. It guides mutational directions toward desired traits by simulating natural selection on wild-type proteins and evaluates the effects of variants based on their fitness to perform the function. We assess the proposed approach using a public database and two new databases for a variety of variant effect prediction tasks, which encompass a diverse set of proteins and assays from different taxa. The prediction results achieve state-of-the-art performance over other zero-shot learning methods for both single-site mutations and deep mutations.
翻译:定向进化在蛋白质工程中扮演着不可或缺的角色,它通过修改现有蛋白质序列以获得新功能或增强功能。准确预测蛋白质变体的效应需要深入理解蛋白质的结构和功能。尽管大型自监督语言模型在仅使用蛋白质序列进行零样本推理时展现了卓越性能,但这些模型本质上无法解释蛋白质结构的空间特征,而后者对于理解蛋白质折叠稳定性和内部分子相互作用至关重要。本文提出了一种新颖的预训练框架,该框架级联了序列与几何分析器,用于处理蛋白质的一级和三级结构。通过模拟野生型蛋白质上的自然选择,它将突变方向引导至所需性状,并根据变体执行功能时的适应度评估其效应。我们使用一个公共数据库和两个新数据库对提出的方法进行了评估,这些数据库涵盖了来自不同分类群的多种蛋白质和实验。在单位点突变和深度突变的预测任务中,该方法相比其他零样本学习方法取得了最先进的性能。