In the realm of future home-assistant robots, 3D articulated object manipulation is essential for enabling robots to interact with their environment. Many existing studies make use of 3D point clouds as the primary input for manipulation policies. However, this approach encounters challenges due to data sparsity and the significant cost associated with acquiring point cloud data, which can limit its practicality. In contrast, RGB images offer high-resolution observations using cost effective devices but lack spatial 3D geometric information. To overcome these limitations, we present a novel image-based robotic manipulation framework. This framework is designed to capture multiple perspectives of the target object and infer depth information to complement its geometry. Initially, the system employs an eye-on-hand RGB camera to capture an overall view of the target object. It predicts the initial depth map and a coarse affordance map. The affordance map indicates actionable areas on the object and serves as a constraint for selecting subsequent viewpoints. Based on the global visual prior, we adaptively identify the optimal next viewpoint for a detailed observation of the potential manipulation success area. We leverage geometric consistency to fuse the views, resulting in a refined depth map and a more precise affordance map for robot manipulation decisions. By comparing with prior works that adopt point clouds or RGB images as inputs, we demonstrate the effectiveness and practicality of our method. In the project webpage (https://sites.google.com/view/imagemanip), real world experiments further highlight the potential of our method for practical deployment.
翻译:在未来家庭辅助机器人领域,三维铰接物体操控对于使机器人能够与环境交互至关重要。许多现有研究采用三维点云作为操控策略的主要输入。然而,这种方法面临数据稀疏性和获取点云数据的高昂成本等挑战,限制了其实用性。相比之下,RGB图像虽能通过低成本设备提供高分辨率观测,但缺乏空间三维几何信息。为克服这些局限,我们提出了一种新颖的基于图像的机器人操控框架。该框架旨在捕获目标物体的多个视角,并推断深度信息以补充其几何结构。系统首先使用眼在手上RGB相机捕获目标物体的整体视图,预测初始深度图和粗略的可操作图。可操作图标示了物体上可执行操作的区域,并作为后续视角选择的约束条件。基于全局视觉先验,我们自适应地确定最佳下一视角,以详细观察潜在的可操控成功区域。我们利用几何一致性融合多视角信息,生成优化的深度图和更精确的可操作图,以支持机器人操控决策。通过与采用点云或RGB图像作为输入的先前工作进行比较,我们展示了该方法的有效性和实用性。在项目网页(https://sites.google.com/view/imagemanip)上,真实世界实验进一步凸显了该方法实际部署的潜力。