The integration of large language models (LLMs) with external tools has significantly expanded the capabilities of AI agents. However, as the diversity of both LLMs and tools increases, selecting the optimal model-tool combination becomes a high-dimensional optimization challenge. Existing approaches often rely on a single model or fixed tool-calling logic, failing to exploit the performance variations across heterogeneous model-tool pairs. In this paper, we present ATLAS (Adaptive Tool-LLM Alignment and Synergistic Invocation), a dual-path framework for dynamic tool usage in cross-domain complex reasoning. ATLAS operates via a dual-path approach: (1) \textbf{training-free cluster-based routing} that exploits empirical priors for domain-specific alignment, and (2) \textbf{RL-based multi-step routing} that explores autonomous trajectories for out-of-distribution generalization. Extensive experiments across 15 benchmarks demonstrate that our method outperforms closed-source models like GPT-4o, surpassing existing routing methods on both in-distribution (+10.1%) and out-of-distribution (+13.1%) tasks. Furthermore, our framework shows significant gains in visual reasoning by orchestrating specialized multi-modal tools.
翻译:大语言模型(LLMs)与外部工具的集成显著扩展了AI Agent的能力边界。然而,随着LLMs和工具多样性的增加,选择最优模型-工具组合成为高维优化难题。现有方法通常依赖单一模型或固定工具调用逻辑,未能充分利用异构模型-工具对之间的性能差异。本文提出ATLAS(自适应工具-大模型对齐与协同调用),一种用于跨领域复杂推理中动态工具使用的双路径框架。ATLAS通过双路径机制运行:(1)**免训练的集群路由**,利用经验先验实现领域特定对齐;(2)**基于强化学习的多步路由**,探索自主轨迹以实现分布外泛化。在15个基准测试上的广泛实验表明,我们的方法在分布内任务(+10.1%)和分布外任务(+13.1%)上均优于GPT-4o等闭源模型,超越现有路由方法。此外,通过编排专用多模态工具,该框架在视觉推理任务上展现出显著性能提升。