Deep neural-network-based language models (LMs) are increasingly applied to large-scale protein sequence data to predict protein function. However, being largely black-box models and thus challenging to interpret, current protein LM approaches do not contribute to a fundamental understanding of sequence-function mappings, hindering rule-based biotherapeutic drug development. We argue that guidance drawn from linguistics, a field specialized in analytical rule extraction from natural language data, can aid with building more interpretable protein LMs that are more likely to learn relevant domain-specific rules. Differences between protein sequence data and linguistic sequence data require the integration of more domain-specific knowledge in protein LMs compared to natural language LMs. Here, we provide a linguistics-based roadmap for protein LM pipeline choices with regard to training data, tokenization, token embedding, sequence embedding, and model interpretation. Incorporating linguistic ideas into protein LMs enables the development of next-generation interpretable machine-learning models with the potential of uncovering the biological mechanisms underlying sequence-function relationships.
翻译:基于深度神经网络的语言模型正越来越多地应用于大规模蛋白质序列数据以预测蛋白质功能。然而,由于目前多数蛋白质语言模型本质上是黑箱模型且难以解释,它们未能促进对序列-功能映射关系的基础性理解,从而阻碍了基于规则的生物治疗药物开发。我们认为,从语言学(一门专长于从自然语言数据中提取分析性规则的学科)中汲取指导,有助于构建更具可解释性的蛋白质语言模型,这类模型更有可能学习到相关的领域特定规则。蛋白质序列数据与语言序列数据之间的差异,要求蛋白质语言模型相比自然语言模型需整合更多领域特异性知识。本文基于语言学原理,针对训练数据、标记化、标记嵌入、序列嵌入及模型解释等蛋白质语言模型构建环节,提供了一条路线图。将语言学思想融入蛋白质语言模型,将推动下一代可解释机器学习模型的发展,有望揭示序列-功能关系背后的生物学机制。