Monocular 3D object detection reveals an economical but challenging task in autonomous driving. Recently center-based monocular methods have developed rapidly with a great trade-off between speed and accuracy, where they usually depend on the object center's depth estimation via 2D features. However, the visual semantic features without sufficient pixel geometry information, may affect the performance of clues for spatial 3D detection tasks. To alleviate this, we propose MonoPGC, a novel end-to-end Monocular 3D object detection framework with rich Pixel Geometry Contexts. We introduce the pixel depth estimation as our auxiliary task and design depth cross-attention pyramid module (DCPM) to inject local and global depth geometry knowledge into visual features. In addition, we present the depth-space-aware transformer (DSAT) to integrate 3D space position and depth-aware features efficiently. Besides, we design a novel depth-gradient positional encoding (DGPE) to bring more distinct pixel geometry contexts into the transformer for better object detection. Extensive experiments demonstrate that our method achieves the state-of-the-art performance on the KITTI dataset.
翻译:单目3D目标检测在自动驾驶领域是一项经济但具有挑战性的任务。近年来,以中心点为基础的单目方法在速度与精度之间取得了良好的平衡,其通常依赖于通过二维特征对目标中心进行深度估计。然而,缺乏足够像素几何信息的视觉语义特征可能会影响空间3D检测任务中线索的性能。为解决这一问题,我们提出了MonoPGC,一种新颖的端到端单目3D目标检测框架,其利用丰富的像素几何上下文。我们引入像素深度估计作为辅助任务,并设计了深度交叉注意力金字塔模块(DCPM),将局部和全局深度几何知识注入视觉特征。此外,我们提出了深度空间感知Transformer(DSAT),以高效整合3D空间位置与深度感知特征。同时,我们设计了一种新颖的深度梯度位置编码(DGPE),为Transformer引入更清晰的像素几何上下文,从而提升目标检测效果。大量实验表明,我们的方法在KITTI数据集上达到了最先进的性能。