Previous evaluations on 6DoF object pose tracking have presented obvious limitations along with the development of this area. In particular, the evaluation protocols are not unified for different methods, the widely-used YCBV dataset contains significant annotation error, and the error metrics also may be biased. As a result, it is hard to fairly compare the methods, which has became a big obstacle for developing new algorithms. In this paper we contribute a unified benchmark to address the above problems. For more accurate annotation of YCBV, we propose a multi-view multi-object global pose refinement method, which can jointly refine the poses of all objects and view cameras, resulting in sub-pixel sub-millimeter alignment errors. The limitations of previous scoring methods and error metrics are analyzed, based on which we introduce our improved evaluation methods. The unified benchmark takes both YCBV and BCOT as base datasets, which are shown to be complementary in scene categories. In experiments, we validate the precision and reliability of the proposed global pose refinement method with a realistic semi-synthesized dataset particularly for YCBV, and then present the benchmark results unifying learning&non-learning and RGB&RGBD methods, with some finds not discovered in previous studies.
翻译:先前关于六自由度物体姿态追踪的评估在该领域发展过程中呈现出明显局限。具体而言,不同方法的评估协议不统一,广泛使用的YCBV数据集存在显著标注误差,且误差度量也可能存在偏差。这使得各方法难以公平比较,已成为新算法开发的重要障碍。本文提出统一基准以解决上述问题。为获得更精确的YCBV标注,我们提出一种多视角多目标全局位姿精化方法,可联合精化所有物体与观察摄像机的位姿,实现亚像素亚毫米级别的对齐误差。通过分析以往评分方法与误差度量的局限性,我们引入改进的评估方法。该统一基准采用YCBV与BCOT作为基础数据集,两者在场景类别上具有互补性。实验中,我们利用专为YCBV设计的半合成真实数据集验证了所提全局位姿精化方法的精度与可靠性,并展示了融合学习/非学习方法及RGB/RGBD方法的基准结果,发现了先前研究未曾揭示的结论。