Protein representation learning has primarily benefited from the remarkable development of language models (LMs). Accordingly, pre-trained protein models also suffer from a problem in LMs: a lack of factual knowledge. The recent solution models the relationships between protein and associated knowledge terms as the knowledge encoding objective. However, it fails to explore the relationships at a more granular level, i.e., the token level. To mitigate this, we propose Knowledge-exploited Auto-encoder for Protein (KeAP), which performs token-level knowledge graph exploration for protein representation learning. In practice, non-masked amino acids iteratively query the associated knowledge tokens to extract and integrate helpful information for restoring masked amino acids via attention. We show that KeAP can consistently outperform the previous counterpart on 9 representative downstream applications, sometimes surpassing it by large margins. These results suggest that KeAP provides an alternative yet effective way to perform knowledge enhanced protein representation learning.
翻译:蛋白质表示学习主要受益于语言模型(LMs)的显著发展。然而,预训练蛋白质模型也面临语言模型中的一个固有问题:缺乏事实性知识。最近的解决方案将蛋白质与相关知识术语之间的关系建模为知识编码目标,但未能从更细粒度的层级(即词元层级)探索这些关系。为缓解这一问题,我们提出了知识利用型蛋白质自编码器(KeAP),该方法在词元层级进行知识图谱探索以进行蛋白质表示学习。实践中,非掩码氨基酸迭代查询相关知识词元,通过注意力机制提取并整合有用信息以恢复掩码氨基酸。实验表明,KeAP在9个代表性下游应用中持续优于此前方法,有时甚至大幅超越。这些结果证明,KeAP为进行知识增强的蛋白质表示学习提供了一种替代且有效的途径。