6D Object Pose Estimation is a crucial yet challenging task in computer vision, suffering from a significant lack of large-scale datasets. This scarcity impedes comprehensive evaluation of model performance, limiting research advancements. Furthermore, the restricted number of available instances or categories curtails its applications. To address these issues, this paper introduces Omni6DPose, a substantial dataset characterized by its diversity in object categories, large scale, and variety in object materials. Omni6DPose is divided into three main components: ROPE (Real 6D Object Pose Estimation Dataset), which includes 332K images annotated with over 1.5M annotations across 581 instances in 149 categories; SOPE(Simulated 6D Object Pose Estimation Dataset), consisting of 475K images created in a mixed reality setting with depth simulation, annotated with over 5M annotations across 4162 instances in the same 149 categories; and the manually aligned real scanned objects used in both ROPE and SOPE. Omni6DPose is inherently challenging due to the substantial variations and ambiguities. To address this challenge, we introduce GenPose++, an enhanced version of the SOTA category-level pose estimation framework, incorporating two pivotal improvements: Semantic-aware feature extraction and Clustering-based aggregation. Moreover, we provide a comprehensive benchmarking analysis to evaluate the performance of previous methods on this large-scale dataset in the realms of 6D object pose estimation and pose tracking.
翻译:6D物体姿态估计是计算机视觉中至关重要且极具挑战性的任务,长期以来面临大规模数据集严重匮乏的问题。这种稀缺性阻碍了对模型性能的全面评估,限制了研究进展。此外,可用实例或类别的数量有限也制约了其应用范围。为应对这些问题,本文提出了Omni6DPose,这是一个在物体类别多样性、规模以及物体材质种类方面均具有显著优势的大型数据集。Omni6DPose主要由三个部分组成:ROPE(真实6D物体姿态估计数据集),包含33.2万张图像,标注超过150万个姿态,涵盖149个类别的581个实例;SOPE(模拟6D物体姿态估计数据集),包含47.5万张在混合现实环境中通过深度模拟生成的图像,标注超过500万个姿态,涵盖相同149个类别的4162个实例;以及用于ROPE和SOPE的、经过手动对齐的真实扫描物体。由于存在显著的类内差异和姿态歧义,Omni6DPose本身具有很高的挑战性。为应对这一挑战,我们提出了GenPose++,这是对当前最优类别级姿态估计框架的增强版本,包含两项关键改进:语义感知特征提取和基于聚类的特征聚合。此外,我们提供了全面的基准测试分析,以评估先前方法在此大规模数据集上,在6D物体姿态估计和姿态跟踪领域的性能表现。