6-DoF robotic grasping is a long-lasting but unsolved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but perform unsatisfactorily on photometrically challenging objects, e.g., objects in transparent or reflective materials. The bottleneck lies in that the surface of these objects can not reflect back accurate depth due to the absorption or refraction of light. In this paper, in contrast to exploiting the inaccurate depth data, we propose the first RGB-only 6-DoF grasping pipeline called MonoGraspNet that utilizes stable 2D features to simultaneously handle arbitrary object grasping and overcome the problems induced by photometrically challenging objects. MonoGraspNet leverages keypoint heatmap and normal map to recover the 6-DoF grasping poses represented by our novel representation parameterized with 2D keypoints with corresponding depth, grasping direction, grasping width, and angle. Extensive experiments in real scenes demonstrate that our method can achieve competitive results in grasping common objects and surpass the depth-based competitor by a large margin in grasping photometrically challenging objects. To further stimulate robotic manipulation research, we additionally annotate and open-source a multi-view and multi-scene real-world grasping dataset, containing 120 objects of mixed photometric complexity with 20M accurate grasping labels.
翻译:六自由度机器人抓取是一个长期存在但尚未解决的问题。现有方法利用强大的三维网络从深度传感器中提取几何抓取表征,在常见物体上展现出卓越精度,但在光度复杂物体(如透明或反光材质物体)上表现欠佳。其瓶颈在于此类物体表面因光线吸收或折射无法准确反射回深度数据。本文提出首个仅依赖RGB图像的六自由度抓取管道MonoGraspNet,该管道利用稳定的二维特征同时处理任意物体抓取,并克服光度复杂物体带来的问题。MonoGraspNet通过关键点热力图和法向量图恢复由我们提出的新型表征参数化的六自由度抓取姿态——该表征以二维关键点及其对应深度、抓取方向、抓取宽度和角度进行参数化。真实场景的大量实验表明,我们的方法在抓取常见物体时能达到竞争力结果,而在抓取光度复杂物体时大幅超越基于深度的方法。为进一步推动机器人操作研究,我们还额外标注并开源了一个多视角多场景的真实世界抓取数据集,包含120个混合光度复杂度的物体,具有2000万个精确抓取标签。