Advances in large language models (LLMs) have recently opened new and promising avenues for small-molecule drug discovery. Yet existing LLM-based approaches for molecular generation often suffer from high rates of invalid and low-quality ligand candidates, a result of the syntactic limitations of current models with regard to molecular strings. In this paper, we introduce $\texttt{ToolMol}$, an evolutionary agentic framework for de novo drug design. $\texttt{ToolMol}$ combines a multi-objective genetic algorithm with an agentic LLM operator that iteratively updates the ligand population. We build a comprehensive toolbox of RDKit-backed functions that allows our agentic operator to consisently make precise ligand modifications. $\texttt{ToolMol}$ achieves state-of-the-art performance on multi-objective property optimization tasks, discovering drug-like and synthesizable ligands that have $>10\%$ stronger predicted binding affinity compared to existing methods, evaluated on three protein targets. $\texttt{ToolMol}$ ligands additionally achieve state-of-the-art results in gold-standard Absolute Binding Free Energy scores, gaining over existing methods by over $35\%$. By studying chain-of-thought reasoning traces, we observe that tool-calling enables the model to more faithfully execute its planned modifications, efficiently exploiting the strong chemical prior knowledge in LLMs.
翻译:摘要:大语言模型(LLMs)的最新进展为小分子药物发现开辟了新的有前景路径。然而,现有基于LLM的分子生成方法常因当前模型在分子字符串语法上的局限性,导致生成大量无效和低质量配体候选物。本文提出$\texttt{ToolMol}$——一种用于从头药物设计的进化智能体框架。$\texttt{ToolMol}$将多目标遗传算法与智能体LLM算子相结合,通过迭代方式更新配体种群。我们构建了基于RDKit功能的综合工具箱,使智能体算子能够持续进行精确的配体修饰。在多目标属性优化任务中,$\texttt{ToolMol}$达到了最先进的性能,针对三个蛋白质靶标的评估表明,其发现的类药且可合成配体的预测结合亲和力比现有方法强$>10\%$。此外,$\texttt{ToolMol}$配体在黄金标准绝对结合自由能评分中取得最先进结果,相比现有方法提升幅度超过$35\%$。通过分析思维链推理过程,我们发现工具调用能力使模型能够更忠实地执行其计划修饰,从而高效利用LLM中蕴含的强大化学先验知识。