Driven by deep learning techniques, perception technology in autonomous driving has developed rapidly in recent years. To achieve accurate and robust perception capabilities, autonomous vehicles are often equipped with multiple sensors, making sensor fusion a crucial part of the perception system. Among these fused sensors, radars and cameras enable a complementary and cost-effective perception of the surrounding environment regardless of lighting and weather conditions. This review aims to provide a comprehensive guideline for radar-camera fusion, particularly concentrating on perception tasks related to object detection and semantic segmentation. Based on the principles of the radar and camera sensors, we delve into the data processing process and representations, followed by an in-depth analysis and summary of radar-camera fusion datasets. In the review of methodologies in radar-camera fusion, we address interrogative questions, including "why to fuse", "what to fuse", "where to fuse", "when to fuse", and "how to fuse", subsequently discussing various challenges and potential research directions within this domain. To ease the retrieval and comparison of datasets and fusion methods, we also provide an interactive website: https://XJTLU-VEC.github.io/Radar-Camera-Fusion.
翻译:受深度学习技术驱动,近年来自动驾驶中的感知技术发展迅速。为实现准确且鲁棒的感知能力,自动驾驶车辆通常配备多种传感器,这使得传感器融合成为感知系统的关键组成部分。在融合传感器中,雷达与相机能够基于互补且经济高效的方式,在不同光照和天气条件下感知周围环境。本综述旨在为雷达-相机融合提供全面指南,特别聚焦于目标检测和语义分割相关感知任务。我们依据雷达和相机传感器原理,深入探讨数据处理流程及表示方式,随后对雷达-相机融合数据集进行深度分析与总结。在雷达-相机融合方法的综述中,我们探讨了包括“为何融合”“融合什么”“何处融合”“何时融合”以及“如何融合”在内的疑问性问题,进而讨论该领域内的多种挑战与潜在研究方向。为方便数据集的检索与融合方法的比较,我们还提供了交互式网站:https://XJTLU-VEC.github.io/Radar-Camera-Fusion。