Generative protein language models are a natural way to design new proteins with desired functions. However, current models are either difficult to direct to produce a protein from a specific family of interest, or must be trained on a large multiple sequence alignment (MSA) from the specific family of interest, making them unable to benefit from transfer learning across families. To address this, we propose $\textbf{P}$r$\textbf{o}$tein $\textbf{E}$volutionary $\textbf{T}$ransformer (PoET), an autoregressive generative model of whole protein families that learns to generate sets of related proteins as sequences-of-sequences across tens of millions of natural protein sequence clusters. PoET can be used as a retrieval-augmented language model to generate and score arbitrary modifications conditioned on any protein family of interest, and can extrapolate from short context lengths to generalize well even for small families. This is enabled by a unique Transformer layer; we model tokens sequentially within sequences while attending between sequences order invariantly, allowing PoET to scale to context lengths beyond those used during training. In extensive experiments on deep mutational scanning datasets, we show that PoET outperforms existing protein language models and evolutionary sequence models for variant function prediction across proteins of all MSA depths. We also demonstrate PoET's ability to controllably generate new protein sequences.
翻译:生成式蛋白质语言模型是设计具有特定功能的新型蛋白质的自然途径。然而,现有模型要么难以定向生成目标蛋白质家族中的蛋白质,要么必须针对特定家族的大型多序列比对(MSA)进行训练,从而无法利用跨家族的迁移学习。为解决这一问题,我们提出蛋白质进化Transformer(PoET)——一种全蛋白质家族的自回归生成模型,该模型通过学习在数千万个天然蛋白质序列簇中将相关蛋白质集合生成为序列之序列。PoET可作为检索增强型语言模型,针对任意感兴趣的蛋白质家族生成和评估任意修饰,并能从短上下文长度进行外推,即使对小型家族也能实现良好的泛化能力。这得益于独特的Transformer层:我们在序列内部按顺序对标记进行建模,同时以顺序不变的方式处理序列间的注意力,使PoET能够扩展到训练时未使用的上下文长度。在基于深度突变扫描数据集的大量实验中,我们证明PoET在跨所有MSA深度的蛋白质变体功能预测任务上优于现有蛋白质语言模型和进化序列模型。我们还展示了PoET可控生成新型蛋白质序列的能力。