Object localization, and more specifically object pose estimation, in large industrial spaces such as warehouses and production facilities, is essential for material flow operations. Traditional approaches rely on artificial artifacts installed in the environment or excessively expensive equipment, that is not suitable at scale. A more practical approach is to utilize existing cameras in such spaces in order to address the underlying pose estimation problem and to localize objects of interest. In order to leverage state-of-the-art methods in deep learning for object pose estimation, large amounts of data need to be collected and annotated. In this work, we provide an approach to the annotation of large datasets of monocular images without the need for manual labor. Our approach localizes cameras in space, unifies their location with a motion capture system, and uses a set of linear mappings to project 3D models of objects of interest at their ground truth 6D pose locations. We test our pipeline on a custom dataset collected from a system of eight cameras in an industrial setting that mimics the intended area of operation. Our approach was able to provide consistent quality annotations for our dataset with 26, 482 object instances at a fraction of the time required by human annotators.
翻译:在仓库和生产设施等大型工业空间中,物体定位(特别是物体姿态估计)对于物流操作至关重要。传统方法依赖于环境中安装的人工标记物或成本过高的设备,难以大规模部署。更实用的方法是利用这些空间中现有的摄像头来解决底层姿态估计问题,并定位目标物体。为利用深度学习领域最先进的方法进行物体姿态估计,需要收集和标注大量数据。本文提出一种无需人工劳动即可标注大规模单目图像数据集的方法。该方法通过空间相机定位、将相机位置与运动捕捉系统统一,并利用一组线性映射将目标物体的三维模型投影至其真实六维姿态位置进行标注。我们在模拟实际作业环境的工业场景中,基于八摄像头系统采集的自定义数据集验证了该管线。实验表明,该方法能够为包含26,482个物体实例的数据集提供质量一致的标注,且标注耗时仅为人工标注的极小部分。