Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning are limited to a predefined set of affordances, thus limiting the adaptability of robots in real-world environments. In this paper, we propose a new method for language-conditioned affordance-pose joint learning in 3D point clouds. Given a 3D point cloud object, our method detects the affordance region and generates appropriate 6-DoF poses for any unconstrained affordance label. Our method consists of an open-vocabulary affordance detection branch and a language-guided diffusion model that generates 6-DoF poses based on the affordance text. We also introduce a new high-quality dataset for the task of language-driven affordance-pose joint learning. Intensive experimental results demonstrate that our proposed method works effectively on a wide range of open-vocabulary affordances and outperforms other baselines by a large margin. In addition, we illustrate the usefulness of our method in real-world robotic applications. Our code and dataset are publicly available at https://3DAPNet.github.io
翻译:可操作区域检测与姿态估计在众多机器人应用中具有重要意义。两者结合可增强机器人的操作能力,其中生成的姿态能够促进对应的可操作任务。以往的可操作区域-姿态联合学习方法局限于预定义的可操作类别,限制了机器人在真实环境中的适应性。本文提出了一种面向语言条件化三维点云中可操作区域-姿态联合学习的新方法。对于给定的三维点云物体,我们的方法能够检测可操作区域,并为任意未约束的可操作标签生成合适的6自由度姿态。该方法包含一个开放词汇可操作区域检测分支,以及一个基于可操作文本生成6自由度姿态的语言引导扩散模型。我们还为语言驱动的可操作区域-姿态联合学习任务引入了一个高质量新数据集。大量实验结果表明,所提方法在各类开放词汇可操作区域上均能有效工作,且显著优于其他基准方法。此外,我们展示了该方法在真实机器人应用中的实用性。代码与数据集已开源于 https://3DAPNet.github.io