We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the scene with a 9-DoF pose transformation. Our method is semi-automatic and works on commonly-available RGB videos, without requiring a depth sensor. Many steps are performed automatically, and the tasks performed by humans are simple, well-specified, and require only limited reasoning in 3D. This makes them feasible for crowd-sourcing and has allowed us to construct a large-scale dataset by annotating real-estate videos from YouTube. Our dataset CAD-Estate offers 108K instances of 12K unique CAD models placed in the 3D representations of 21K videos. In comparison to Scan2CAD, the largest existing dataset with CAD model annotations on real scenes, CAD-Estate has 8x more instances and 4x more unique CAD models. We showcase the benefits of pre-training a Mask2CAD model on CAD-Estate for the task of automatic 3D object reconstruction and pose estimation, demonstrating that it leads to improvements on the popular Scan2CAD benchmark. We will release the data by mid July 2023.
翻译:摘要:我们提出了一种方法,用于对包含多个物体的复杂场景视频进行标注,为每个物体生成全局一致的3D表示。我们从数据库中选取CAD模型标注每个物体,并通过9自由度位姿变换将其置于场景的3D坐标系中。该方法为半自动方式,仅需常规RGB视频输入,无需深度传感器。其中多数步骤自动完成,人工操作简单明确,仅需有限的3D空间推理能力。这使得该流程适用于众包,并使我们能够通过标注YouTube上的房地产视频构建大规模数据集。我们的CAD-Estate数据集包含21K视频的3D表示中放置的12K个独特CAD模型,共计108K实例。与现有最大的真实场景CAD模型标注数据集Scan2CAD相比,CAD-Estate的实例数量是其8倍,独特CAD模型数量是其4倍。我们展示了在CAD-Estate上预训练Mask2CAD模型对自动3D物体重建与位姿估计任务的提升效果,验证了其在主流Scan2CAD基准上的性能改进。该数据集将于2023年7月中旬公开。