Monocular 3D object detection plays a crucial role in autonomous driving. However, existing monocular 3D detection algorithms depend on 3D labels derived from LiDAR measurements, which are costly to acquire for new datasets and challenging to deploy in novel environments. Specifically, this study investigates the pipeline for training a monocular 3D object detection model on a diverse collection of 3D and 2D datasets. The proposed framework comprises three components: (1) a robust monocular 3D model capable of functioning across various camera settings, (2) a selective-training strategy to accommodate datasets with differing class annotations, and (3) a pseudo 3D training approach using 2D labels to enhance detection performance in scenes containing only 2D labels. With this framework, we could train models on a joint set of various open 3D/2D datasets to obtain models with significantly stronger generalization capability and enhanced performance on new dataset with only 2D labels. We conduct extensive experiments on KITTI/nuScenes/ONCE/Cityscapes/BDD100K datasets to demonstrate the scaling ability of the proposed method.
翻译:单目3D目标检测在自动驾驶中扮演着关键角色。然而,现有单目3D检测算法依赖基于激光雷达测量获得的3D标签,这类标签对新数据集而言获取成本高昂,且难以部署至新环境。本研究系统探究了在多样化3D与2D数据集集合上训练单目3D目标检测模型的流程。提出的框架包含三个组件:(1)能够在不同相机设置下正常工作的鲁棒单目3D模型,(2)适应不同类别标注数据集的择性训练策略,以及(3)利用2D标签提升仅含2D标签场景检测性能的伪3D训练方法。通过该框架,我们可在多种开放3D/2D数据集的联合集合上训练模型,从而获得具有显著更强泛化能力且在新2D标签数据集上性能增强的模型。我们在KITTI/nuScenes/ONCE/Cityscapes/BDD100K数据集上开展大量实验,验证了所提方法的规模化能力。