6-DoF object-agnostic grasping in unstructured environments is a critical yet challenging task in robotics. Most current works use non-optimized approaches to sample grasp locations and learn spatial features without concerning the grasping task. This paper proposes GraNet, a graph-based grasp pose generation framework that translates a point cloud scene into multi-level graphs and propagates features through graph neural networks. By building graphs at the scene level, object level, and grasp point level, GraNet enhances feature embedding at multiple scales while progressively converging to the ideal grasping locations by learning. Our pipeline can thus characterize the spatial distribution of grasps in cluttered scenes, leading to a higher rate of effective grasping. Furthermore, we enhance the representation ability of scalable graph networks by a structure-aware attention mechanism to exploit local relations in graphs. Our method achieves state-of-the-art performance on the large-scale GraspNet-1Billion benchmark, especially in grasping unseen objects (+11.62 AP). The real robot experiment shows a high success rate in grasping scattered objects, verifying the effectiveness of the proposed approach in unstructured environments.
翻译:六自由度通用物体抓取在非结构化环境中是机器人学中一项关键但具有挑战性的任务。当前多数工作采用非优化方法采样抓取位置,并在不涉及抓取任务的前提下学习空间特征。本文提出GraNet——一种基于图的抓取姿态生成框架,该框架将点云场景转化为多层级图,并通过图神经网络传播特征。通过在场景级、物体级和抓取点级构建图结构,GraNet在多重尺度上增强特征嵌入,同时通过学习逐步收敛至理想抓取位置。因此,我们的流程能够表征杂乱场景中抓取姿态的空间分布,从而提高有效抓取的比率。此外,我们通过一种结构感知注意力机制增强可扩展图网络的表示能力,以捕捉图中的局部关系。本文方法在大型GraspNet-1Billion基准测试中实现最先进性能,尤其在抓取未见物体方面取得显著提升(AP提升11.62%)。真实机器人实验表明,该方法在抓取散乱物体时具有高成功率,验证了所提方法在非结构化环境中的有效性。