Robotic systems that assist humans should be capable of adapting their behaviors to individual user preferences. For instance, users may want a robot arm to adjust the amount of force it applies while folding their laundry or cleaning furniture. Natural language provides an intuitive way for humans to communicate such preferences. Recent progress in language-conditioned robot policies has shown that robots can successfully use language prompts to determine what task to perform. However, extending the same approach to realize how the task should be performed requires detailed labels describing the preferences or styles of trajectories in the task data. Not only is collecting such annotations challenging, but conditioning directly on these labels may also fail to provide fine-grained control over a continuous range of behaviors. For example, it can be difficult to convey the exact force that a robot must apply through abstract instructions like "apply a bit more pressure than before". Therefore, in this work, we propose using language to reason over preferred behaviors instead of directly generating them. We first learn a structured latent representation that organizes user preferences according to differences in the corresponding trajectories. Then, given a preference prompt, we use a foundation model to interpret this latent space and choose a value that produces the desired behavior. Through both simulation and real-world experiments, we show that selecting robot behaviors from an intuitively structured latent space enables more precise adaptation to user preferences while requiring significantly fewer preference labels than language-conditioned policies.


翻译:摘要:辅助人类的机器人系统应能根据个体用户偏好调整其行为。例如,用户可能希望机器臂在折叠衣物或清洁家具时调整施加的力。自然语言为人类传达此类偏好提供了直观方式。语言条件机器人策略的最新进展表明,机器人能成功利用语言提示确定待执行任务。然而,将相同方法扩展至实现任务执行方式,需要任务数据中描述轨迹偏好或风格的详细标注。此类标注的收集不仅具有挑战性,且直接以这些标签为条件可能无法对连续行为范围提供细粒度控制。例如,通过"比之前稍微多施加一点压力"这类抽象指令传递机器人需施加的准确力度存在困难。因此,本研究提出利用语言推理偏好行为而非直接生成行为。我们首先学习一种结构化潜在表征,根据相应轨迹差异对用户偏好进行组织。随后,给定偏好提示时,利用基础模型解释该潜在空间,并选择能产生所需行为的对应值。通过仿真与真实实验,我们证明:从直观结构化的潜在空间中选择机器人行为,能实现对用户偏好的更精准适应,同时所需偏好标注量显著少于语言条件策略。

0
下载
关闭预览

相关内容

【伯克利博士论文】将机器人的表征与人类对齐
专知会员服务
46+阅读 · 2023年8月27日
《结合机器人行为以实现安全、智能的执行》
专知会员服务
17+阅读 · 2023年7月4日
《机器人语言》美陆军5年项目46页技术总结报告,2023年
专知会员服务
41+阅读 · 2023年5月17日
【机器人】机器人PID控制
产业智能官
10+阅读 · 2018年11月25日
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员