While multiagent systems have shown promise for tackling complex tasks via specialization, finetuning multiple agents simultaneously faces two key challenges: (1) credit assignment across agents, and (2) sample efficiency of expensive multiagent rollouts. In this work, we propose finetuning multiagent systems with per-action process rewards from AI feedback (MAPPA) to address both. Through assigning credit to individual agent actions rather than only at task completion, MAPPA enables fine-grained supervision without ground truth labels while extracting maximal training signal from each rollout. We demonstrate our approach on competition math problems and tool-augmented data analysis tasks. On unseen math problems, MAPPA achieves +5.0--17.5pp on AIME and +7.8--17.2pp on AMC. For data analysis tasks, our method improves success rate by +12.5pp while quality metrics improve by up to 30%, validating that per-action supervision can lead to improvements across different multiagent system on various domains. By addressing these challenges, our work takes a first step toward scaling multiagent systems for complex, long-horizon tasks with minimal human supervision.


翻译:尽管多智能体系统通过专业化分工在处理复杂任务方面展现出潜力,但同时微调多个智能体面临两个关键挑战:(1) 跨智能体的信用分配问题,(2) 昂贵多智能体推演的样本效率问题。本研究提出通过人工智能反馈的逐动作过程奖励(MAPPA)来微调多智能体系统,以同时解决这两个挑战。MAPPA通过将信用分配给单个智能体动作而非仅在任务完成时分配,实现了无需真实标签的细粒度监督,同时从每次推演中提取最大化的训练信号。我们在数学竞赛题和工具增强的数据分析任务上验证了该方法。在未见过的数学问题上,MAPPA在AIME上实现了+5.0-17.5个百分点的提升,在AMC上实现了+7.8-17.2个百分点的提升。对于数据分析任务,我们的方法将成功率提高了+12.5个百分点,质量指标提升高达30%,验证了逐动作监督能够推动不同领域多智能体系统的全面改进。通过解决这些挑战,我们的工作为在最小化人工监督条件下扩展多智能体系统处理复杂长程任务迈出了第一步。

0
下载
关闭预览

相关内容

《多智能体大语言模型系统的可靠决策研究》
专知会员服务
43+阅读 · 2月2日
迈向智能体系统规模化的科学
专知会员服务
23+阅读 · 2025年12月12日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
43+阅读 · 2025年4月18日
面向大模型多智能体系统的多维评估方法
专知会员服务
36+阅读 · 2025年4月15日
【AAMAS教程】多智能体优化,241页ppt
专知会员服务
68+阅读 · 2024年3月1日
基于多智能体深度强化学习的体系任务分配方法
专知会员服务
160+阅读 · 2023年5月4日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
面向多智能体博弈对抗的对手建模框架
专知
19+阅读 · 2022年9月28日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
18+阅读 · 2020年9月9日
经典书《斯坦福大学-多智能体系统》532页pdf
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2月9日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
10+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
9+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
6+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
12+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
15+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
6+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
13+阅读 · 9月21日
相关VIP内容
《多智能体大语言模型系统的可靠决策研究》
专知会员服务
43+阅读 · 2月2日
迈向智能体系统规模化的科学
专知会员服务
23+阅读 · 2025年12月12日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
43+阅读 · 2025年4月18日
面向大模型多智能体系统的多维评估方法
专知会员服务
36+阅读 · 2025年4月15日
【AAMAS教程】多智能体优化,241页ppt
专知会员服务
68+阅读 · 2024年3月1日
基于多智能体深度强化学习的体系任务分配方法
专知会员服务
160+阅读 · 2023年5月4日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
19+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员