Molecule generation with desired properties has grown immensely in popularity by disruptively changing the way scientists design molecular structures and providing support for chemical and materials design. However, despite the promising outcome, previous machine learning-based deep generative models suffer from a reliance on complex, task-specific fine-tuning, limited dimensional latent spaces, or the quality of expert rules. In this work, we propose MolGen, a pre-trained molecular language model that effectively learns and shares knowledge across multiple generation tasks and domains. Specifically, we pre-train MolGen with the chemical language SELFIES on more than 100 million unlabelled molecules. We further propose multi-task molecular prefix tuning across several molecular generation tasks and different molecular domains (synthetic & natural products) with a self-feedback mechanism. Extensive experiments show that MolGen can obtain superior performances on well-known molecular generation benchmark datasets. The further analysis illustrates that MolGen can accurately capture the distribution of molecules, implicitly learn their structural characteristics, and efficiently explore the chemical space with the guidance of multi-task molecular prefix tuning. Codes, datasets, and the pre-trained model will be available in https://github.com/zjunlp/MolGen.
翻译:具有期望性质的分子生成因颠覆性地改变了科学家设计分子结构的方式,并为化学与材料设计提供支持而日益普及。然而,尽管取得令人鼓舞的成果,以往基于机器学习的深度生成模型仍存在局限性,包括依赖复杂的任务专用微调、潜在空间维度受限或专家规则质量参差不齐等问题。本文提出预训练分子语言模型MolGen,该模型能有效学习并跨多个生成任务与领域共享知识。具体而言,我们采用化学语言SELFIES在超过1亿个无标签分子上对MolGen进行预训练,并进一步提出多任务分子前缀调优方法,通过自反馈机制在多个分子生成任务及不同分子领域(合成产物与天然产物)间实现协同优化。大量实验表明,MolGen在公认的分子生成基准数据集上可获得优异性能。进一步分析揭示,MolGen能精准捕捉分子分布、隐式学习其结构特征,并在多任务分子前缀调优引导下高效探索化学空间。相关代码、数据集及预训练模型将发布于https://github.com/zjunlp/MolGen。