In this paper, we propose a novel training strategy called SupFusion, which provides an auxiliary feature level supervision for effective LiDAR-Camera fusion and significantly boosts detection performance. Our strategy involves a data enhancement method named Polar Sampling, which densifies sparse objects and trains an assistant model to generate high-quality features as the supervision. These features are then used to train the LiDAR-Camera fusion model, where the fusion feature is optimized to simulate the generated high-quality features. Furthermore, we propose a simple yet effective deep fusion module, which contiguously gains superior performance compared with previous fusion methods with SupFusion strategy. In such a manner, our proposal shares the following advantages. Firstly, SupFusion introduces auxiliary feature-level supervision which could boost LiDAR-Camera detection performance without introducing extra inference costs. Secondly, the proposed deep fusion could continuously improve the detector's abilities. Our proposed SupFusion and deep fusion module is plug-and-play, we make extensive experiments to demonstrate its effectiveness. Specifically, we gain around 2% 3D mAP improvements on KITTI benchmark based on multiple LiDAR-Camera 3D detectors.
翻译:本文提出了一种名为SupFusion的新型训练策略,该策略通过提供辅助特征级监督实现高效的激光雷达-相机融合,显著提升了检测性能。我们的策略包含一种名为极坐标采样的数据增强方法,该方法可对稀疏目标进行稠密化处理,并训练辅助模型生成高质量特征作为监督信号。随后利用这些特征训练激光雷达-相机融合模型,使融合特征优化以模拟生成的高质量特征。此外,我们提出了一种简单而有效的深度融合模块,结合SupFusion策略后,其性能持续优于现有融合方法。基于此,本方案具备以下优势:第一,SupFusion引入的辅助特征级监督可在不增加推理成本的前提下提升激光雷达-相机检测性能;第二,所提深度融合模块能够持续增强检测器能力。所提出的SupFusion与深度融合模块具有即插即用特性,通过大量实验验证了其有效性。具体而言,在KITTI基准测试上,基于多种激光雷达-相机3D检测器,我们实现了约2%的3D平均精度提升。