Machine learning-based design has gained traction in the sciences, most notably in the design of small molecules, materials, and proteins, with societal implications spanning drug development and manufacturing, plastic degradation, and carbon sequestration. When designing objects to achieve novel property values with machine learning, one faces a fundamental challenge: how to push past the frontier of current knowledge, distilled from the training data into the model, in a manner that rationally controls the risk of failure. If one trusts learned models too much in extrapolation, one is likely to design rubbish. In contrast, if one does not extrapolate, one cannot find novelty. Herein, we ponder how one might strike a useful balance between these two extremes. We focus in particular on designing proteins with novel property values, although much of our discussion addresses machine learning-based design more broadly.
翻译:基于机器学习的设计在科学领域日益受到重视,尤其是在小分子、材料和蛋白质设计方面,其对药物开发与制造、塑料降解以及碳封存等社会应用产生深远影响。当利用机器学习设计具有新颖性能值的对象时,人们面临一个根本性挑战:如何在合理控制失败风险的前提下,突破从训练数据中提炼并嵌入模型的现有知识边界。若在推断过程中过度信任学习模型,极有可能设计出无价值的产物;反之,若不进行推断,则无法发现新颖性。本文旨在探求如何在这两个极端之间取得平衡。尽管我们的讨论广泛涉及基于机器学习的设计,但重点关注如何设计具有新颖性能值的蛋白质。