End-to-end learning for visual robotic manipulation is known to suffer from sample inefficiency, requiring large numbers of demonstrations. The spatial roto-translation equivariance, or the SE(3)-equivariance can be exploited to improve the sample efficiency for learning robotic manipulation. In this paper, we present SE(3)-equivariant models for visual robotic manipulation from point clouds that can be trained fully end-to-end. By utilizing the representation theory of the Lie group, we construct novel SE(3)-equivariant energy-based models that allow highly sample efficient end-to-end learning. We show that our models can learn from scratch without prior knowledge and yet are highly sample efficient (5~10 demonstrations are enough). Furthermore, we show that our models can generalize to tasks with (i) previously unseen target object poses, (ii) previously unseen target object instances of the category, and (iii) previously unseen visual distractors. We experiment with 6-DoF robotic manipulation tasks to validate our models' sample efficiency and generalizability. Codes are available at: https://github.com/tomato1mule/edf
翻译:端到端学习的视觉机器人操控已知存在样本效率低下的问题,需要大量的演示数据。利用空间旋转-平移等变性,即SE(3)-等变性,可提升机器人操控学习的样本效率。本文提出了基于点云的视觉机器人操控SE(3)-等变模型,这些模型能够实现完全端到端的训练。通过利用李群的表示理论,我们构建了新型SE(3)-等变能量模型,该模型可实现高度样本高效的端到端学习。我们证明,该模型无需先验知识即可从零开始学习,且样本效率极高(仅需5~10次演示)。此外,我们展示了该模型能够泛化至以下任务场景:(i) 未见过目标物体姿态;(ii) 未见过类别内的目标物体实例;(iii) 未见过的视觉干扰物。我们通过6自由度机器人操控任务实验,验证了模型的样本效率与泛化能力。代码已开源:https://github.com/tomato1mule/edf