A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast MCMC sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.
翻译:机器学习驱动的蛋白质工程长期目标是加速发现改善已知蛋白质功能的新型突变。我们提出一种用于计算机内蛋白质进化的采样框架,该框架支持混合搭配各类无监督模型(如蛋白质语言模型)以及从序列预测蛋白质功能的有监督模型。通过组合这些模型,我们旨在提升评估未见突变的能力,并将搜索约束于可能包含功能性蛋白质的序列空间区域。该框架无需任何模型微调或重新训练,直接在离散蛋白质空间中构建专家乘积分布。不同于经典定向进化中采用的暴力搜索或随机采样,我们引入一种基于梯度提出有利突变的快速MCMC采样器。我们在广泛适应性景观上及多种不同预训练无监督模型(包括650M参数蛋白质语言模型)上进行了计算机内定向进化实验。结果表明,该方法能高效发现具有高进化似然性及距野生型蛋白质多个突变位点的预测活性变体,证明我们的采样器为基于机器学习的蛋白质工程提供了实用且有效的新范式。