Previous evaluations on 6DoF object pose tracking have presented obvious limitations along with the development of this area. In particular, the evaluation protocols are not unified for different methods, the widely-used YCBV dataset contains significant annotation error, and the error metrics also may be biased. As a result, it is hard to fairly compare the methods, which has became a big obstacle for developing new algorithms. In this paper we contribute a unified benchmark to address the above problems. For more accurate annotation of YCBV, we propose a multi-view multi-object global pose refinement method, which can jointly refine the poses of all objects and view cameras, resulting in sub-pixel sub-millimeter alignment errors. The limitations of previous scoring methods and error metrics are analyzed, based on which we introduce our improved evaluation methods. The unified benchmark takes both YCBV and BCOT as base datasets, which are shown to be complementary in scene categories. In experiments, we validate the precision and reliability of the proposed global pose refinement method with a realistic semi-synthesized dataset particularly for YCBV, and then present the benchmark results unifying learning&non-learning and RGB&RGBD methods, with some finds not discovered in previous studies.
翻译:先前关于六自由度物体姿态追踪的评估方法随着该领域的发展暴露出明显局限。具体而言,不同方法的评估协议尚未统一,广泛使用的YCBV数据集存在显著标注误差,且误差度量指标可能存在偏差。因此,方法间的公平比较变得困难,这已成为新算法开发的重要障碍。本文构建了统一基准数据集以解决上述问题。针对YCBV标注精度不足,我们提出多视角多目标全局姿态优化方法,可联合优化所有物体与视角相机的姿态,实现亚像素级对齐误差。通过分析现有评分方法与误差度量指标的局限性,我们引入改进的评估方案。该统一基准采用YCBV与BCOT作为基础数据集,两者在场景类别上具有互补性。实验部分,我们首先利用为YCBV特制的真实感半合成数据集验证了所提全局姿态优化方法的精度与可靠性,继而呈现了融合学习方法/非学习方法及RGB/RGBD方法的基准测试结果,发现了前人研究未曾揭示的现象。