Multi-Agent Reinforcement Learning (MARL) provides a powerful framework for learning coordination in multi-agent systems. However, applying MARL to robotics remains challenging due to their high-dimensional continuous joint action spaces, complex reward design, and non-stationarity from concurrently learning agents. On the other hand, humans often learn complex coordination with the help of coaches, who guide learning through carefully designed curricula and detailed feedback. Building on the reasoning capabilities of foundation models, we argue that these models can similarly coach robots to learn coordination. Motivated by this, we propose CRAFT: Coaching Reinforcement learning Autonomously using Foundation models for learning coordination Tasks, a framework that leverages foundation models to act as a "coach" for multi-robot coordination. CRAFT automatically decomposes long-horizon coordination tasks into sequences of subtasks using the planning capability of Large Language Models (LLMs). Then, CRAFT trains each subtask using LLM-generated reward functions, and refines them through a Vision Language Model (VLM)-guided reward-refinement loop. We evaluate CRAFT on multi-quadruped navigation and bimanual manipulation tasks, and demonstrate its capability to learn complex coordination behaviors. In addition, in a multi-quadruped navigation setting, we show that our learned policies transfer to the real world. Project website is https://iconlab.negarmehr.com/CRAFT/


翻译:摘要:多智能体强化学习(MARL)为多智能体系统中的协调学习提供了强大的框架。然而,由于机器人领域存在高维连续联合动作空间、复杂的奖励设计以及并发学习智能体导致的非平稳性,将MARL应用于机器人仍具挑战性。另一方面,人类往往借助教练的指导来学习复杂的协调技能——教练通过精心设计的课程和详细反馈引导学习过程。基于基础模型的推理能力,我们认为这些模型可以类似地引导机器人学习协调技能。受此启发,我们提出CRAFT(基于基础模型自主进行协调任务强化学习辅导)框架,利用基础模型充当多机器人协调的“教练”。CRAFT首先利用大语言模型(LLMs)的规划能力,将长时域协调任务自动分解为子任务序列。随后,通过LLM生成的奖励函数训练每个子任务,并借助视觉语言模型(VLM)引导的奖励精化循环对奖励进行优化。我们在多四足机器人导航和双臂操作任务上评估了CRAFT,证明了其学习复杂协调行为的能力。此外,在多四足机器人导航场景中,我们展示了学得的策略可迁移至真实世界。项目网站:https://iconlab.negarmehr.com/CRAFT/

0
下载
关闭预览

相关内容

开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
《空战战术多智能体强化学习中的可解释性》最新报告
专知会员服务
86+阅读 · 2024年10月25日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
《多智能体强化学习:基础与现代方法》2023最新320页书稿
专知会员服务
130+阅读 · 2023年10月26日
「博弈论视角下多智能体强化学习」研究综述
专知会员服务
185+阅读 · 2022年4月30日
「基于通信的多智能体强化学习」 进展综述
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员