The complexity of teaching humanoid robots new tasks is one of the major reasons hindering their widespread adoption in the industry. While Imitation Learning (IL), particularly Action Chunking with Transformers (ACT), enables rapid task acquisition, there is no consensus yet on the optimal sensory hardware required for manipulation tasks. This paper benchmarks 14 sensor combinations on the Unitree G1 humanoid robot equipped with three-finger hands for two manipulation tasks. We explicitly evaluate the integration of tactile and proprioceptive modalities alongside active vision. Our analysis demonstrates that strategic sensor selection can outperform complex configurations in data-limited regimes while reducing computational overhead. We develop an open-source Unified Ablation Framework that utilizes sensor masking on a comprehensive master dataset. Results indicate that additional modalities often degrade performance for IL with limited data. A minimal active stereo-camera setup outperformed complex multi-sensor configurations, achieving 87.5% success in a spatial generalization task and 94.4% in a structured manipulation task. Conversely, adding pressure sensors to this setup reduced success to 67.3% in the latter task due to a low signal-to-noise ratio. We conclude that in data-limited regimes, active vision offers a superior trade-off between robustness and complexity. While tactile modalities may require larger datasets to be effective, our findings validate that strategic sensor selection is critical for designing an efficient learning process.


翻译:模仿学习(Imitation Learning,IL)尤其是基于Transformer的行动分块(Action Chunking with Transformers,ACT)能够实现快速任务获取,但关于操作任务所需的最佳感知硬件尚未达成共识。本文在配备三指手的Unitree G1人形机器人上针对两个操作任务对14种传感器组合进行了基准测试。我们明确评估了触觉与本体感觉模态与主动视觉的集成效果。分析表明,在数据受限条件下,战略性传感器选择能够优于复杂配置,同时减少计算开销。我们开发了一个开源统一消融框架,通过对综合主数据集进行传感器掩码处理。结果表明,在有限数据下,额外模态通常会降低模仿学习性能。一种极简的主动双目摄像头配置优于复杂的多传感器配置,在空间泛化任务中达到87.5%的成功率,在结构化操作任务中达到94.4%的成功率。相反,由于信噪比低,向该配置添加压力传感器后,后一任务的成功率降至67.3%。我们得出结论:在数据受限条件下,主动视觉在鲁棒性和复杂性之间提供了更优的权衡。尽管触觉模态可能需要更大数据集才能发挥作用,但我们的发现验证了战略性的传感器选择对于设计高效学习过程至关重要。

0
下载
关闭预览

相关内容

深度学习时代的模仿学习:新型分类体系与最新研究进展
机器人中的深度生成模型:多模态演示学习的综述
专知会员服务
40+阅读 · 2024年8月9日
【NeurIPS 2020】生成对抗性模仿学习的f-Divergence
专知会员服务
26+阅读 · 2020年10月9日
浅谈主动学习(Active Learning)
凡人机器学习
32+阅读 · 2020年6月18日
原创 | Attention Modeling for Targeted Sentiment
黑龙江大学自然语言处理实验室
25+阅读 · 2017年11月5日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
7+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
相关VIP内容
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员