Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We propose Team-of-Thoughts, a heterogeneous MAS framework that treats diverse models as specialized tools within an orchestrator-driven paradigm. Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration, which identifies models with superior coordination and synthesis capabilities, and (2) Agent Self-Assessment, a protocol where tool agents profile their own domain-specific strengths to guide selection. At inference, the orchestrator dynamically activates the most compatible agents based on these profiles to maximize capability coverage. Across five mathematical reasoning and code generation benchmarks, Team-of-Thoughts consistently outperforms individual models and existing MAS baselines. Notably, on AIME24 and LiveCodeBench, Team-of-Thoughts achieves 96.00% and 77.91% accuracy, respectively, significantly improving over homogeneous role-play baselines (80.00% and 65.93%).
翻译:现有的大多数多智能体系统(MAS)通常依赖同构模型配置,未能充分利用不同后训练架构中蕴含的多样化专业知识。我们提出“团队思考”(Team-of-Thoughts)这一异构MAS框架,将多样化的模型视为编排器驱动范式中的专用工具。团队思考引入了两个新颖组件:(1)编排器校准,用于识别具有卓越协调与综合能力的模型;(2)智能体自我评估,即一种协议,工具智能体在此协议中剖析自身领域特定优势以指导选择。在推理时,编排器基于这些评估动态激活最兼容的智能体,以最大化能力覆盖范围。在五个数学推理与代码生成基准测试上,团队思考持续优于单个模型及现有MAS基线。值得注意的是,在AIME24和LiveCodeBench上,团队思考分别达到了96.00%和77.91%的准确率,相比同构角色扮演基线(80.00%和65.93%)有显著提升。