Antibodies are vital proteins offering robust protection for the human body from pathogens. The development of general protein and antibody-specific pre-trained language models both facilitate antibody prediction tasks. However, there have been limited studies that comprehensively explore the representation capability of distinct pre-trained language models on different antibody tasks. To investigate the problem, we aim to answer several key questions in this paper, such as how pre-trained language models perform in antibody tasks with different specificity and how introducing specific biological mechanisms to the pre-training process can benefit the model. Additionally, we evaluate if the learned antibody pre-trained representations can be applied to real-world antibody problems, like drug discovery and immune process understanding. Previously, no benchmark available largely hindered the study to answer these questions. To aid in our investigation, we provide an AnTibody Understanding Evaluation (ATUE) benchmark. We comprehensively evaluate the performance of protein pre-trained language models by empirical study along with conclusions and new insights. Our ATUE and code are released at https://github.com/dqwang122/EATLM.
翻译:抗体是重要的蛋白质,能够为人体提供强有力的病原体保护。通用蛋白质及抗体特异性预训练语言模型的发展均促进了抗体预测任务。然而,目前尚缺乏系统探究不同预训练语言模型在不同抗体任务中表征能力的研究。为深入探讨该问题,本文旨在回答若干关键问题,例如:预训练语言模型在特异性不同的抗体任务中表现如何,以及在预训练过程中引入特定生物学机制能否使模型受益。此外,我们评估了学习到的抗体预训练表征能否应用于真实世界的抗体问题,如药物发现和免疫过程理解。此前,缺乏基准测试在很大程度上阻碍了这些问题的研究。为此,我们提供了抗体理解评估基准(ATUE)。我们通过实证研究全面评估了蛋白质预训练语言模型的性能,并得出相关结论与新见解。ATUE及代码已发布在https://github.com/dqwang122/EATLM。