We demonstrate how the often overlooked inherent properties of large-scale LiDAR point clouds can be effectively utilized for self-supervised representation learning. In pursuit of this goal, we design a highly data-efficient feature pre-training backbone that considerably reduces the need for tedious 3D annotations to train state-of-the-art object detectors. We propose Masked AutoEncoder for LiDAR point clouds (MAELi) that intuitively leverages the sparsity of LiDAR point clouds in both the encoder and decoder during reconstruction. Our approach results in more expressive and useful features, which can be directly applied to downstream perception tasks, such as 3D object detection for autonomous driving. In a novel reconstruction schema, MAELi distinguishes between free and occluded space and employs a new masking strategy that targets the LiDAR's inherent spherical projection. To demonstrate the potential of MAELi, we pre-train one of the most widely-used 3D backbones in an end-to-end manner and show the effectiveness of our unsupervised pre-trained features on various 3D object detection architectures. Our method achieves significant performance improvements when only a small fraction of labeled frames is available for fine-tuning object detectors. For instance, with ~800 labeled frames, MAELi features enhance a SECOND model by +10.79APH/LEVEL 2 on Waymo Vehicles.
翻译:我们展示了如何有效利用大规模激光雷达点云中常被忽视的固有属性进行自监督表示学习。为此,我们设计了一种高数据效率的特征预训练主干网络,显著减少训练最先进目标检测器所需的大量3D标注工作。我们提出面向激光雷达点云的掩码自编码器(MAELi),在编码器和解码器的重建过程中直观地利用了激光雷达点云的稀疏性。该方法能生成更具表现力和实用性的特征,可直接应用于自动驾驶等下游感知任务(如3D目标检测)。在一种新型重建方案中,MAELi区分自由空间与遮挡空间,并采用针对激光雷达固有球面投影的掩码策略。为展示MAELi的潜力,我们以端到端方式预训练了最广泛使用的3D主干网络之一,并在多种3D目标检测架构上验证了无监督预训练特征的有效性。当仅有少量标注帧可用于目标检测器微调时,该方法能实现显著的性能提升。例如,在Waymo车辆数据集上,使用约800个标注帧时,MAELi特征使SECOND模型的APH/LEVEL 2指标提升+10.79。