Currently, object detection applications in construction are almost based on pure 2D data (both image and annotation are 2D-based), resulting in the developed artificial intelligence (AI) applications only applicable to some scenarios that only require 2D information. However, most advanced applications usually require AI agents to perceive 3D spatial information, which limits the further development of the current computer vision (CV) in construction. The lack of 3D annotated datasets for construction object detection worsens the situation. Therefore, this study creates and releases a virtual dataset with 3D annotations named VCVW-3D, which covers 15 construction scenes and involves ten categories of construction vehicles and workers. The VCVW-3D dataset is characterized by multi-scene, multi-category, multi-randomness, multi-viewpoint, multi-annotation, and binocular vision. Several typical 2D and monocular 3D object detection models are then trained and evaluated on the VCVW-3D dataset to provide a benchmark for subsequent research. The VCVW-3D is expected to bring considerable economic benefits and practical significance by reducing the costs of data construction, prototype development, and exploration of space-awareness applications, thus promoting the development of CV in construction, especially those of 3D applications.
翻译:目前,建筑施工领域的目标检测应用几乎完全基于二维数据(图像和标注均为二维化),导致所开发的人工智能应用仅适用于仅需二维信息的场景。然而,大多数高级应用通常要求智能体感知三维空间信息,这制约了当前计算机视觉技术在施工领域的进一步发展。施工目标检测数据集缺乏三维标注的现状加剧了这一困境。为此,本研究创建并发布了名为VCVW-3D的虚拟三维标注数据集,该数据集涵盖15个施工场景及10类施工车辆与工人。VCVW-3D数据集具有多场景、多类别、多随机性、多视角、多标注与双目视觉的特点。我们随后在VCVW-3D数据集上训练并评估了多种典型二维目标检测模型与单目三维目标检测模型,为后续研究提供了基准。通过降低数据构建、原型开发及空间感知应用探索的成本,VCVW-3D有望带来显著的经济效益与实践价值,从而推动计算机视觉在施工领域的发展,特别是三维应用方面。