Active perception in vision-based robotic manipulation aims to move the camera toward more informative observation viewpoints, thereby providing high-quality perceptual inputs for downstream tasks. Most existing active perception methods rely on iterative optimization, leading to high time and motion costs, and are tightly coupled with task-specific objectives, which limits their transferability. In this paper, we propose a general one-shot multimodal active perception framework for robotic manipulation. The framework enables direct inference of optimal viewpoints and comprises a data collection pipeline and an optimal viewpoint prediction network. Specifically, the framework decouples viewpoint quality evaluation from the overall architecture, supporting heterogeneous task requirements. Optimal viewpoints are defined through systematic sampling and evaluation of candidate viewpoints, after which large-scale training datasets are constructed via domain randomization. Moreover, a multimodal optimal viewpoint prediction network is developed, leveraging cross-attention to align and fuse multimodal features and directly predict camera pose adjustments. The proposed framework is instantiated in robotic grasping under viewpoint-constrained environments. Experimental results demonstrate that active perception guided by the framework significantly improves grasp success rates. Notably, real-world evaluations achieve nearly double the grasp success rate and enable seamless sim-to-real transfer without additional fine-tuning, demonstrating the effectiveness of the proposed framework.


翻译:基于视觉的机器人操作中的主动感知旨在将相机移动到更具信息量的观测视角,从而为下游任务提供高质量的感知输入。现有的大多数主动感知方法依赖于迭代优化,导致较高的时间和运动成本,并且与特定任务目标紧密耦合,这限制了其可迁移性。在本文中,我们提出了一种面向机器人操作的单次多模态主动感知通用框架。该框架能够直接推断最优观测视角,并包含一个数据收集流程和一个最优视角预测网络。具体而言,该框架将视角质量评估从整体架构中解耦,以支持异构任务需求。最优视角通过对候选视角进行系统采样和评估来定义,随后通过领域随机化构建大规模训练数据集。此外,我们开发了一个多模态最优视角预测网络,利用交叉注意力机制对齐和融合多模态特征,并直接预测相机位姿调整。所提出的框架在视角受限环境下的机器人抓取任务中进行了实例化。实验结果表明,由该框架引导的主动感知显著提高了抓取成功率。值得注意的是,真实世界评估实现了接近翻倍的抓取成功率,并且无需额外微调即可实现从仿真到现实的无缝迁移,这证明了所提出框架的有效性。

0
下载
关闭预览

相关内容

【干货书】基于深度学习的机器人感知与认知,638页pdf
专知会员服务
114+阅读 · 2022年7月29日
这可能是「多模态机器学习」最通俗易懂的介绍
计算机视觉life
113+阅读 · 2018年12月20日
【机器视觉】机器视觉全面解析
产业智能官
12+阅读 · 2018年11月12日
机器学习必知的15大框架
云栖社区
16+阅读 · 2017年12月10日
报名 | 让机器读懂你的意图——人体姿态估计入门
人工智能头条
10+阅读 · 2017年9月19日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
面向国防作战的最佳自主与蜂群无人机技术
专知会员服务
7+阅读 · 7月28日
《异构人类团队的协作决策过程混合建模研究》
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
5+阅读 · 7月27日
美空军新型反无人机部队初探
专知会员服务
9+阅读 · 7月27日
《防空交战流程的概率建模研究》
专知会员服务
12+阅读 · 7月27日
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
7+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
10+阅读 · 7月26日
相关VIP内容
【干货书】基于深度学习的机器人感知与认知,638页pdf
专知会员服务
114+阅读 · 2022年7月29日
相关基金
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员