The generation of molecules with desired properties has gained tremendous popularity, revolutionizing the way scientists design molecular structures and providing valuable support for chemical and drug design. However, despite the potential of language models in molecule generation, they face numerous challenges such as the generation of syntactically or chemically flawed molecules, narrow domain focus, and limitations in creating diverse and directionally feasible molecules due to a dearth of annotated data or external molecular databases. To this end, we introduce MolGen, a pre-trained molecular language model tailored specifically for molecule generation. MolGen acquires intrinsic structural and grammatical insights by reconstructing over 100 million molecular SELFIES, while facilitating knowledge transfer between different domains through domain-agnostic molecular prefix tuning. Moreover, we present a self-feedback paradigm that inspires the pre-trained model to align with the ultimate goal of producing molecules with desirable properties. Extensive experiments demonstrate that MolGen achieves superior performance on well-known molecule generation benchmarks. Further analysis shows that MolGen can accurately capture molecule distributions, implicitly learn their structural characteristics, and efficiently explore chemical space. The pre-trained model, codes, and datasets are publicly available for future research at https://github.com/zjunlp/MolGen.
翻译:具有期望特性的分子生成技术已获得广泛关注,这一技术革新了科学家设计分子结构的方式,并为化学与药物设计提供了重要支持。然而,尽管语言模型在分子生成领域展现出潜力,仍面临诸多挑战:生成语法或化学结构存在缺陷的分子、领域聚焦狭窄、因标注数据或外部分子数据库匮乏而难以产生多样且方向可行的分子。为此,我们提出了MolGen——一种专为分子生成设计的预训练分子语言模型。该模型通过重构超过1亿条分子SELFIES序列获取内在结构与语法知识,同时采用领域无关的分子前缀调优策略促进跨领域知识迁移。此外,我们提出了一种自反馈范式,引导预训练模型与生成理想特性分子的终极目标保持一致。大量实验表明,MolGen在公认的分子生成基准测试中取得了优越性能。进一步分析显示,MolGen能精准捕捉分子分布特征、隐式学习分子结构特性,并高效探索化学空间。预训练模型、代码及数据集已在https://github.com/zjunlp/MolGen 开源供后续研究使用。