In recent years, modern techniques in deep learning and large-scale datasets have led to impressive progress in 3D instance segmentation, grasp pose estimation, and robotics. This allows for accurate detection directly in 3D scenes, object- and environment-aware grasp prediction, as well as robust and repeatable robotic manipulation. This work aims to integrate these recent methods into a comprehensive framework for robotic interaction and manipulation in human-centric environments. Specifically, we leverage 3D reconstructions from a commodity 3D scanner for open-vocabulary instance segmentation, alongside grasp pose estimation, to demonstrate dynamic picking of objects, and opening of drawers. We show the performance and robustness of our model in two sets of real-world experiments including dynamic object retrieval and drawer opening, reporting a 51% and 82% success rate respectively. Code of our framework as well as videos are available on: https://spot-compose.github.io/.
翻译:摘要:近年来,深度学习技术与大规模数据集的现代方法在三维实例分割、抓取姿态估计及机器人学领域取得了显著进展。这使得直接在三维场景中进行精确检测、实现物体与环境感知的抓取预测,以及执行稳健且可重复的机器人操作成为可能。本研究旨在将这些前沿方法整合为一个面向人类环境中的机器人交互与操作的综合性框架。具体而言,我们利用商用三维扫描仪重建的三维数据进行开放词汇实例分割,并结合抓取姿态估计,演示了物体的动态拾取与抽屉的开启操作。我们通过两组真实世界实验(包括动态物体检索与抽屉开启)展示了模型的性能与鲁棒性,分别报告了51%与82%的成功率。本框架的代码及演示视频可在以下链接获取:https://spot-compose.github.io/。