Large language models (LLMs) have shown remarkable potential in various domains, but they often lack the ability to access and reason over domain-specific knowledge and tools. In this paper, we introduced CACTUS (Chemistry Agent Connecting Tool-Usage to Science), an LLM-based agent that integrates cheminformatics tools to enable advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama2-7b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b and Mistral-7b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with domain-specific tools, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment. Furthermore, CACTUS represents a significant milestone in the field of cheminformatics, offering an adaptable tool for researchers engaged in chemistry and molecular discovery. By integrating the strengths of open-source LLMs with domain-specific tools, CACTUS has the potential to accelerate scientific advancement and unlock new frontiers in the exploration of novel, effective, and safe therapeutic candidates, catalysts, and materials. Moreover, CACTUS's ability to integrate with automated experimentation platforms and make data-driven decisions in real time opens up new possibilities for autonomous discovery.
翻译:摘要:大型语言模型(LLMs)在多个领域展现出显著潜力,但往往缺乏获取并推理领域特定知识与工具的能力。本文提出CACTUS(化学智能体连接工具使用与科学),一种基于LLMs的智能体,通过整合化学信息学工具实现化学与分子发现领域的深度推理与问题求解。我们采用包含数千道化学问题的基准测试,基于Gemma-7b、Falcon-7b、MPT-7b、Llama2-7b和Mistral-7b等多样化开源LLMs评估CACTUS性能。结果表明,CACTUS显著优于基础LLMs,其中Gemma-7b与Mistral-7b模型在不同提示策略下均达到最高准确率。此外,我们探究了领域特异性提示与硬件配置对模型性能的影响,凸显提示工程的重要性以及在小规模消费级硬件上部署较小模型的可能性(且不显著损失精度)。通过将开源LLMs的认知能力与领域工具相结合,CACTUS可辅助研究人员完成分子性质预测、相似性搜索及类药性评估等任务。更进一步,CACTUS代表着化学信息学领域的重要里程碑,为从事化学与分子发现的研究人员提供了一种自适应工具。通过整合开源LLMs与领域专用工具的优势,CACTUS有望加速科学进步,为探索新型、高效且安全的候选治疗药物、催化剂及材料开辟新前沿。尤为重要的是,CACTUS具备与自动化实验平台集成、实时数据驱动决策的能力,为自主发现开辟了新途径。