Across scientific domains, generating new models or optimizing existing ones while meeting specific criteria is crucial. Traditional machine learning frameworks for guided design use a generative model and a surrogate model (discriminator), requiring large datasets. However, real-world scientific applications often have limited data and complex landscapes, making data-hungry models inefficient or impractical. We propose a new framework, PropEn, inspired by ``matching'', which enables implicit guidance without training a discriminator. By matching each sample with a similar one that has a better property value, we create a larger training dataset that inherently indicates the direction of improvement. Matching, combined with an encoder-decoder architecture, forms a domain-agnostic generative framework for property enhancement. We show that training with a matched dataset approximates the gradient of the property of interest while remaining within the data distribution, allowing efficient design optimization. Extensive evaluations in toy problems and scientific applications, such as therapeutic protein design and airfoil optimization, demonstrate PropEn's advantages over common baselines. Notably, the protein design results are validated with wet lab experiments, confirming the competitiveness and effectiveness of our approach.
翻译:在众多科学领域中,根据特定标准生成新模型或优化现有模型至关重要。传统的引导设计机器学习框架通常使用生成模型和代理模型(判别器),这需要大量数据集。然而,现实世界的科学应用往往数据有限且问题空间复杂,使得依赖海量数据的模型效率低下或不切实际。我们提出了一种受“匹配”思想启发的新框架PropEn,它无需训练判别器即可实现隐式引导。通过将每个样本与一个具有更优属性值的相似样本进行匹配,我们创建了一个更大的训练数据集,该数据集本质上指明了改进方向。匹配机制与编码器-解码器架构相结合,形成了一个与领域无关的属性增强生成框架。我们证明,使用匹配数据集进行训练可以近似目标属性的梯度,同时保持在数据分布范围内,从而实现高效的设计优化。在玩具问题和科学应用(如治疗性蛋白质设计和翼型优化)中的广泛评估表明,PropEn相较于常见基线方法具有显著优势。值得注意的是,蛋白质设计结果已通过湿实验室实验验证,证实了我们方法的竞争力和有效性。