In this paper, we present GP3D, a novel network for generalized pose estimation in 3D point clouds. The method generalizes to new objects by using both the scene point cloud and the object point cloud with keypoint indexes as input. The network is trained to match the object keypoints to scene points. To address the pose estimation of novel objects we also present a new approach for training pose estimation. The typical solution is a single model trained for pose estimation of a specific object in any scenario. This has several drawbacks: training a model for each object is time-consuming, energy consuming, and by excluding the scenario information the task becomes more difficult. In this paper, we present the opposite solution; a scenario-specific pose estimation method for novel objects that do not require retraining. The network is trained on 1500 objects and is able to learn a generalized solution. We demonstrate that the network is able to correctly predict novel objects, and demonstrate the ability of the network to perform outside of the trained class. We believe that the demonstrated method is a valuable solution for many real-world scenarios. Code and trained network will be made available after publication.
翻译:本文提出GP3D,一种用于三维点云中广义姿态估计的新型网络。该方法通过同时输入场景点云和带有关键点索引的目标点云,实现了对新物体的泛化能力。该网络经过训练,能够将目标关键点与场景点进行匹配。为了解决新物体的姿态估计问题,我们还提出了一种新的姿态估计训练方法。典型的解决方案是针对特定物体在任何场景下进行姿态估计训练单一模型。这种方法存在若干缺陷:为每个物体训练模型耗时耗能,且排除场景信息后任务难度更大。本文提出相反的解决方案:一种无需重新训练即可针对新物体进行场景特定姿态估计的方法。该网络在1500个物体上进行训练,能够学习到通用解决方案。实验证明该网络能正确预测新物体,并展现出在训练类别之外进行推断的能力。我们相信,所提出的方法对许多实际场景具有重要应用价值。代码和训练好的网络将在论文发表后公开。