Autonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird's-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors.
翻译:自动驾驶需要包括3D物体检测、跟踪和分割在内的精确且快速的3D感知系统。尽管近期基于低成本相机的方法展现出良好效果,但易受光线不足或恶劣天气影响,且存在较大定位误差。因此,融合能够提供精确远程测量并在所有环境下稳定工作的低成本雷达与相机虽前景广阔,但尚未得到充分研究。本文提出相机雷达网络(CRN),一种新颖的相机-雷达融合框架,可生成语义丰富且空间精确的鸟瞰图(BEV)特征图以支持多种任务。为克服图像空间信息不足的问题,我们借助稀疏但精确的雷达点,将透视图像特征转换至BEV空间。进一步采用针对多模态空间错位设计的多模态可变形注意力机制,在BEV中聚合图像与雷达特征图。实时配置下的CRN能以20 FPS运行,在nuScenes数据集上性能与激光雷达检测器相当,甚至在100米远距离场景表现更优。此外,离线配置的CRN在nuScenes测试集上取得62.4% NDS、57.5% mAP,在所有基于相机及相机-雷达的3D物体检测器中排名第一。