Large Language Models (LLMs) have revolutionized the field of natural language processing, but they fall short in comprehending biological sequences such as proteins. To address this challenge, we propose InstructProtein, an innovative LLM that possesses bidirectional generation capabilities in both human and protein languages: (i) taking a protein sequence as input to predict its textual function description and (ii) using natural language to prompt protein sequence generation. To achieve this, we first pre-train an LLM on both protein and natural language corpora, enabling it to comprehend individual languages. Then supervised instruction tuning is employed to facilitate the alignment of these two distinct languages. Herein, we introduce a knowledge graph-based instruction generation framework to construct a high-quality instruction dataset, addressing annotation imbalance and instruction deficits in existing protein-text corpus. In particular, the instructions inherit the structural relations between proteins and function annotations in knowledge graphs, which empowers our model to engage in the causal modeling of protein functions, akin to the chain-of-thought processes in natural languages. Extensive experiments on bidirectional protein-text generation tasks show that InstructProtein outperforms state-of-the-art LLMs by large margins. Moreover, InstructProtein serves as a pioneering step towards text-based protein function prediction and sequence design, effectively bridging the gap between protein and human language understanding.
翻译:大型语言模型(LLMs)已彻底改变自然语言处理领域,但在理解蛋白质等生物序列方面仍存在不足。为应对这一挑战,我们提出InstructProtein——一种创新性LLM,具备人类语言与蛋白质语言的双向生成能力:(i)输入蛋白质序列以预测其文本功能描述,(ii)利用自然语言提示蛋白质序列生成。为实现这一目标,我们首先在蛋白质和自然语言语料库上预训练LLM,使其能理解各自语言;随后采用监督式指令微调促进两种语言的语义对齐。为此,我们引入基于知识图谱的指令生成框架,构建高质量指令数据集,以解决现有蛋白质文本语料库中的标注不平衡与指令缺失问题。特别地,指令继承了知识图谱中蛋白质与功能注释之间的结构关系,使模型能够对蛋白质功能进行因果建模,类似于自然语言中的思维链过程。双向蛋白质文本生成任务的大量实验表明,InstructProtein在性能上大幅超越现有最先进的LLMs。此外,InstructProtein作为基于文本的蛋白质功能预测与序列设计的先驱性工作,有效弥合了蛋白质理解与人类语言理解之间的鸿沟。