Architectures that first convert point clouds to a grid representation and then apply convolutional neural networks achieve good performance for radar-based object detection. However, the transfer from irregular point cloud data to a dense grid structure is often associated with a loss of information, due to the discretization and aggregation of points. In this paper, we propose a novel architecture, multi-scale KPPillarsBEV, that aims to mitigate the negative effects of grid rendering. Specifically, we propose a novel grid rendering method, KPBEV, which leverages the descriptive power of kernel point convolutions to improve the encoding of local point cloud contexts during grid rendering. In addition, we propose a general multi-scale grid rendering formulation to incorporate multi-scale feature maps into convolutional backbones of detection networks with arbitrary grid rendering methods. We perform extensive experiments on the nuScenes dataset and evaluate the methods in terms of detection performance and computational complexity. The proposed multi-scale KPPillarsBEV architecture outperforms the baseline by 5.37% and the previous state of the art by 2.88% in Car AP4.0 (average precision for a matching threshold of 4 meters) on the nuScenes validation set. Moreover, the proposed single-scale KPBEV grid rendering improves the Car AP4.0 by 2.90% over the baseline while maintaining the same inference speed.
翻译:摘要:将点云首先转换为网格表示,然后应用卷积神经网络的架构在基于雷达的目标检测中取得了良好性能。然而,从非结构化点云数据到密集网格结构的转换,由于点的离散化和聚合,常伴随信息损失。本文提出了一种新型架构——多尺度KPPillarsBEV,旨在减轻网格渲染的负面影响。具体地,我们提出了一种新型网格渲染方法KPBEV,利用核点卷积的描述能力,在网格渲染过程中改进局部点云上下文的编码。此外,我们提出了一种通用的多尺度网格渲染公式,将多尺度特征图集成到任意网格渲染方法的检测网络卷积骨干中。我们在nuScenes数据集上进行了广泛实验,并从检测性能和计算复杂度两个方面评估了方法。所提出的多尺度KPPillarsBEV架构在nuScenes验证集上的Car AP4.0(匹配阈值为4米的平均精度)方面,比基线提升了5.37%,比先前最先进方法提升了2.88%。此外,所提出的单尺度KPBEV网格渲染在保持相同推理速度的同时,将Car AP4.0相比基线提升了2.90%。