In this survey, we first introduce the background of popular sensors used for self-driving, their data properties, and the corresponding object detection algorithms. Next, we discuss existing datasets that can be used for evaluating multi-modal 3D object detection algorithms. Then we present a review of multi-modal fusion based 3D detection networks, taking a close look at their fusion stage, fusion input and fusion granularity, and how these design choices evolve with time and technology. After the review, we discuss open challenges as well as possible solutions. We hope that this survey can help researchers to get familiar with the field and embark on investigations in the area of multi-modal 3D object detection.
翻译:本综述首先介绍了自动驾驶中常用传感器的背景、其数据特性以及相应的目标检测算法。接着,我们讨论了可用于评估多模态3D目标检测算法的现有数据集。随后,我们对基于多模态融合的3D检测网络进行了回顾,深入分析了它们的融合阶段、融合输入和融合粒度,以及这些设计选择如何随着时间和技术的发展而演变。在回顾之后,我们探讨了当前的开放挑战以及可能的解决方案。希望本综述能够帮助研究人员熟悉该领域,并投身于多模态3D目标检测的研究之中。