This paper addresses the limitations of current datasets for 3D vision tasks in terms of accuracy, size, realism, and suitable imaging modalities for photometrically challenging objects. We propose a novel annotation and acquisition pipeline that enhances existing 3D perception and 6D object pose datasets. Our approach integrates robotic forward-kinematics, external infrared trackers, and improved calibration and annotation procedures. We present a multi-modal sensor rig, mounted on a robotic end-effector, and demonstrate how it is integrated into the creation of highly accurate datasets. Additionally, we introduce a freehand procedure for wider viewpoint coverage. Both approaches yield high-quality 3D data with accurate object and camera pose annotations. Our methods overcome the limitations of existing datasets and provide valuable resources for 3D vision research.
翻译:本文针对光度学挑战性物体在三维视觉任务中现有数据集在精度、规模、真实感及适用成像模态方面的局限性,提出了一种新颖的标注与采集流程以增强现有三维感知和六自由度物体位姿数据集。该方法融入了机器人正向运动学、外部红外追踪器以及改进的标定与标注流程。我们展示了一套安装于机器人末端执行器的多模态传感器平台,并证实其如何被整合到高精度数据集的创建过程中。此外,还提出了一种自由手持操作流程以实现更广泛的视角覆盖。两种方法均能生成具有精确物体和相机位姿标注的高质量三维数据。我们的技术克服了现有数据集的局限性,为三维视觉研究提供了宝贵资源。