The combination of LiDAR and camera modalities is proven to be necessary and typical for 3D object detection according to recent studies. Existing fusion strategies tend to overly rely on the LiDAR modal in essence, which exploits the abundant semantics from the camera sensor insufficiently. However, existing methods cannot rely on information from other modalities because the corruption of LiDAR features results in a large domain gap. Following this, we propose CrossFusion, a more robust and noise-resistant scheme that makes full use of the camera and LiDAR features with the designed cross-modal complementation strategy. Extensive experiments we conducted show that our method not only outperforms the state-of-the-art methods under the setting without introducing an extra depth estimation network but also demonstrates our model's noise resistance without re-training for the specific malfunction scenarios by increasing 5.2\% mAP and 2.4\% NDS.
翻译:根据近期研究,激光雷达与相机模态的结合被证明是三维目标检测中必要且典型的技术路径。现有融合策略本质上过度依赖激光雷达模态,未能充分挖掘相机传感器的丰富语义信息。然而当激光雷达特征受损时,由于产生较大域差异,现有方法无法有效依赖其他模态信息。为此,我们提出CrossFusion——一种更具鲁棒性和抗噪声能力的方案,通过设计的跨模态互补策略充分利用相机与激光雷达特征。大量实验表明,在不引入额外深度估计网络的设置下,我们的方法不仅优于现有最先进方法,更在无需针对特定故障场景重训练的情况下,通过提升5.2%的平均精度均值(mAP)和2.4%的归一化检测分数(NDS),证明了模型优异的抗噪声性能。