Object detection has been extensively utilized in autonomous systems in recent years, encompassing both 2D and 3D object detection. Recent research in this field has primarily centered around multimodal approaches for addressing this issue.In this paper, a multimodal fusion approach based on result feature-level fusion is proposed. This method utilizes the outcome features generated from single modality sources, and fuses them for downstream tasks.Based on this method, a new post-fusing network is proposed for multimodal object detection, which leverages the single modality outcomes as features. The proposed approach, called Multi-Modal Detector based on Result features (MMDR), is designed to work for both 2D and 3D object detection tasks. Compared to previous multimodal models, the proposed approach in this paper performs feature fusion at a later stage, enabling better representation of the deep-level features of single modality sources. Additionally, the MMDR model incorporates shallow global features during the feature fusion stage, endowing the model with the ability to perceive background information and the overall input, thereby avoiding issues such as missed detections.
翻译:目标检测在近年来的自主系统中得到了广泛应用,涵盖2D与3D目标检测。该领域的最新研究主要聚焦于基于多模态的方法来解决该问题。本文提出了一种基于结果特征级融合的多模态融合方法,该方法利用单一模态源生成的结果特征,并将其融合用于下游任务。基于该方法,本文提出了一种用于多模态目标检测的新型后融合网络,该网络将单一模态结果作为特征。所提出的方法称为基于结果特征的多模态检测器(MMDR),旨在同时适用于2D和3D目标检测任务。与以往的多模态模型相比,本文方法在更晚的阶段进行特征融合,从而能够更好地表征单一模态源的深层特征。此外,MMDR模型在特征融合阶段融入了浅层全局特征,使模型具备感知背景信息与整体输入的能力,从而避免漏检等问题。