Current protein language models (PLMs) learn protein representations mainly based on their sequences, thereby well capturing co-evolutionary information, but they are unable to explicitly acquire protein functions, which is the end goal of protein representation learning. Fortunately, for many proteins, their textual property descriptions are available, where their various functions are also described. Motivated by this fact, we first build the ProtDescribe dataset to augment protein sequences with text descriptions of their functions and other important properties. Based on this dataset, we propose the ProtST framework to enhance Protein Sequence pre-training and understanding by biomedical Texts. During pre-training, we design three types of tasks, i.e., unimodal mask prediction, multimodal representation alignment and multimodal mask prediction, to enhance a PLM with protein property information with different granularities and, at the same time, preserve the PLM's original representation power. On downstream tasks, ProtST enables both supervised learning and zero-shot prediction. We verify the superiority of ProtST-induced PLMs over previous ones on diverse representation learning benchmarks. Under the zero-shot setting, we show the effectiveness of ProtST on zero-shot protein classification, and ProtST also enables functional protein retrieval from a large-scale database without any function annotation.
翻译:当前蛋白质语言模型主要通过序列学习蛋白质表示,能够较好地捕捉共进化信息,但无法显式获取蛋白质功能——而功能正是蛋白质表示学习的最终目标。幸运的是,许多蛋白质已具备文本属性描述,其中涵盖了其各类功能信息。基于此观察,我们首先构建了ProtDescribe数据集,为蛋白质序列补充功能及其他重要属性的文本描述。基于该数据集,我们提出ProtST框架,旨在通过生物医学文本增强蛋白质序列预训练与理解。在预训练阶段,我们设计了三种任务类型(单模态掩码预测、多模态表示对齐及多模态掩码预测),以不同粒度向蛋白质语言模型注入蛋白质属性信息,同时保持模型原有的表示能力。在下游任务中,ProtST支持监督学习与零样本预测两种范式。我们在多个表示学习基准上验证了ProtST诱导的蛋白质语言模型相较先前模型的优越性。在零样本设置下,我们证明了ProtST在零样本蛋白质分类中的有效性,且该框架无需任何功能标注即可实现大规模数据库中的功能蛋白质检索。