An effective framework for learning 3D representations for perception tasks is distilling rich self-supervised image features via contrastive learning. However, image-to point representation learning for autonomous driving datasets faces two main challenges: 1) the abundance of self-similarity, which results in the contrastive losses pushing away semantically similar point and image regions and thus disturbing the local semantic structure of the learned representations, and 2) severe class imbalance as pretraining gets dominated by over-represented classes. We propose to alleviate the self-similarity problem through a novel semantically tolerant image-to-point contrastive loss that takes into consideration the semantic distance between positive and negative image regions to minimize contrasting semantically similar point and image regions. Additionally, we address class imbalance by designing a class-agnostic balanced loss that approximates the degree of class imbalance through an aggregate sample-to-samples semantic similarity measure. We demonstrate that our semantically-tolerant contrastive loss with class balancing improves state-of-the art 2D-to-3D representation learning in all evaluation settings on 3D semantic segmentation. Our method consistently outperforms state-of-the-art 2D-to-3D representation learning frameworks across a wide range of 2D self-supervised pretrained models.
翻译:针对感知任务学习三维表示的有效框架是通过对比学习蒸馏丰富的自监督图像特征。然而,自动驾驶数据集中的图像到点云表示学习面临两大挑战:1)大量自相似性导致对比损失将语义相似的点云与图像区域推开,从而破坏所学表示的局部语义结构;2)预训练被过度表示的类别主导引发的严重类别不平衡问题。我们提出通过新型语义容忍图像到点云对比损失来缓解自相似性问题,该损失考虑了正负图像区域之间的语义距离,以最小化语义相似点云与图像区域的对比。此外,我们设计了一种类别无关的平衡损失来解决类别不平衡问题,该损失通过聚合的样本间语义相似度度量来近似类别不平衡程度。实验表明,在3D语义分割的所有评估设置中,我们的具有类别平衡的语义容忍对比损失改进了最先进的2D到3D表示学习。在广泛使用的2D自监督预训练模型上,我们的方法始终优于最先进的2D到3D表示学习框架。