Estimating the 6-DoF pose of a rigid object from a single RGB image is a crucial yet challenging task. Recent studies have shown the great potential of dense correspondence-based solutions, yet improvements are still needed to reach practical deployment. In this paper, we propose a novel pose estimation algorithm named CheckerPose, which improves on three main aspects. Firstly, CheckerPose densely samples 3D keypoints from the surface of the 3D object and finds their 2D correspondences progressively in the 2D image. Compared to previous solutions that conduct dense sampling in the image space, our strategy enables the correspondence searching in a 2D grid (i.e., pixel coordinate). Secondly, for our 3D-to-2D correspondence, we design a compact binary code representation for 2D image locations. This representation not only allows for progressive correspondence refinement but also converts the correspondence regression to a more efficient classification problem. Thirdly, we adopt a graph neural network to explicitly model the interactions among the sampled 3D keypoints, further boosting the reliability and accuracy of the correspondences. Together, these novel components make our CheckerPose a strong pose estimation algorithm. When evaluated on the popular Linemod, Linemod-O, and YCB-V object pose estimation benchmarks, CheckerPose clearly boosts the accuracy of correspondence-based methods and achieves state-of-the-art performances.
翻译:单张RGB图像中刚体物体的6自由度姿态估计是一项关键但具有挑战性的任务。近期研究表明,基于密集对应关系的解决方案潜力巨大,但在实际部署中仍需进一步提升。本文提出一种名为CheckerPose的新型姿态估计算法,其在三个主要方面进行了改进。首先,CheckerPose从3D物体表面密集采样3D关键点,并在2D图像中渐进式地寻找其对应的2D点。与以往在图像空间进行密集采样的方案相比,我们的策略能够在2D网格(即像素坐标)中搜索对应关系。其次,针对3D到2D对应关系,我们为2D图像位置设计了一种紧凑的二进制编码表示。这种表示不仅支持渐进式对应关系精化,还将对应关系回归转化为更高效的分类问题。再次,我们采用图神经网络显式建模采样3D关键点之间的相互作用,进一步提升了对应关系的可靠性和准确性。这些创新组件共同使CheckerPose成为强大的姿态估计算法。在流行的Linemod、Linemod-O和YCB-V物体姿态估计基准测试上,CheckerPose显著提升了基于对应关系方法的精度,并取得了最先进的性能。