Accurate obstacle identification represents a fundamental challenge within the scope of near-field perception for autonomous driving. Conventionally, fisheye cameras are frequently employed for comprehensive surround-view perception, including rear-view obstacle localization. However, the performance of such cameras can significantly deteriorate in low-light conditions, during nighttime, or when subjected to intense sun glare. Conversely, cost-effective sensors like ultrasonic sensors remain largely unaffected under these conditions. Therefore, we present, to our knowledge, the first end-to-end multimodal fusion model tailored for efficient obstacle perception in a bird's-eye-view (BEV) perspective, utilizing fisheye cameras and ultrasonic sensors. Initially, ResNeXt-50 is employed as a set of unimodal encoders to extract features specific to each modality. Subsequently, the feature space associated with the visible spectrum undergoes transformation into BEV. The fusion of these two modalities is facilitated via concatenation. At the same time, the ultrasonic spectrum-based unimodal feature maps pass through content-aware dilated convolution, applied to mitigate the sensor misalignment between two sensors in the fused feature space. Finally, the fused features are utilized by a two-stage semantic occupancy decoder to generate grid-wise predictions for precise obstacle perception. We conduct a systematic investigation to determine the optimal strategy for multimodal fusion of both sensors. We provide insights into our dataset creation procedures, annotation guidelines, and perform a thorough data analysis to ensure adequate coverage of all scenarios. When applied to our dataset, the experimental results underscore the robustness and effectiveness of our proposed multimodal fusion approach.
翻译:精准的障碍物识别是自动驾驶近场感知领域的一个基础性挑战。传统上,鱼眼相机常被用于环视感知,包括后视障碍物定位。然而,在弱光条件、夜间或强烈阳光直射下,此类相机的性能会显著下降。相比之下,超声波传感器等低成本传感器在上述条件下基本不受影响。因此,我们提出了据我们所知首个面向鸟瞰视角(BEV)高效障碍物感知的端到端多模态融合模型,该模型采用鱼眼相机与超声波传感器。首先,采用ResNeXt-50作为一组单模态编码器,提取各模态特定特征。随后,可见光谱相关的特征空间被转换为BEV表示。两种模态的融合通过特征拼接实现。同时,基于超声波光谱的单模态特征图经过内容感知膨胀卷积处理,以缓解融合特征空间中两类传感器之间的对准偏差。最后,融合特征被输入至两阶段语义占用解码器,生成网格级预测以实现精准障碍物感知。我们系统研究了两种传感器多模态融合的最优策略,详细描述了数据集创建流程与标注规范,并通过充分的数据分析确保场景覆盖的完整性。实验结果表明,所提多模态融合方法在数据集上展现了鲁棒性与有效性。