Computational drug discovery, particularly the complex workflows of drug molecule screening and optimization, requires orchestrating dozens of specialized tools in multi-step workflows, yet current AI agents struggle to maintain robust performance and consistently underperform in these high-complexity scenarios. Here we present MolClaw, an autonomous agent that leads drug molecule evaluation, screening, and optimization. It unifies over 30 specialized domain resources through a three-tier hierarchical skill architecture (70 skills in total) that facilitates agent long-term interaction at runtime: tool-level skills standardize atomic operations, workflow-level skills compose them into validated pipelines with quality check and reflection, and a discipline-level skill supplies scientific principles governing planning and verification across all scenarios in the field. Additionally, we introduce MolBench, a benchmark comprising molecular screening, optimization, and end-to-end discovery challenges spanning 8 to 50+ sequential tool calls. MolClaw achieves state-of-the-art performance across all metrics, and ablation studies confirm that gains concentrate on tasks that demand structured workflows while vanishing on those solvable with ad hoc scripting, establishing workflow orchestration competence as the primary capability bottleneck for AI-driven drug discovery.
翻译:计算药物发现,特别是药物分子筛选与优化的复杂工作流,需要协调数十种专业工具完成多步骤工作流程,然而当前的人工智能自主体难以维持稳健性能,在此类高复杂度场景中始终表现欠佳。本文提出MolClaw,这是一个引领药物分子评估、筛选与优化任务的自主体。它通过三层层次化技能架构(共计70项技能)整合了超过30个专业领域资源,支撑自主体在运行时实现长期交互:工具级技能标准化原子操作,工作流级技能将其组合为经过质量检查与反思验证的流水线,学科级技能则为整个领域所有场景下的规划与验证提供科学原理。此外,我们引入了MolBench基准测试,包含涵盖8至50余次连续工具调用的分子筛选、优化及端到端发现挑战。MolClaw在所有评估指标上均取得最先进性能,消融研究证实性能提升主要集中在需要结构化工作流的任务上,而在可通过临时脚本解决的场景中优势消失,这确立了工作流编排能力作为AI驱动药物发现主要能力瓶颈的地位。