We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of free-form language queries. This is a challenging problem because of the 2D-3D ambiguity and the open-vocabulary nature of the target tasks, where obtaining annotated training data in 3D is difficult. The contributions of this work are three-fold. First, we design a new model architecture for open-vocabulary 3D semantic occupancy prediction. The architecture consists of a 2D-3D encoder together with occupancy prediction and 3D-language heads. The output is a dense voxel map of 3D grounded language embeddings enabling a range of open-vocabulary tasks. Second, we develop a tri-modal self-supervised learning algorithm that leverages three modalities: (i) images, (ii) language and (iii) LiDAR point clouds, and enables training the proposed architecture using a strong pre-trained vision-language model without the need for any 3D manual language annotations. Finally, we demonstrate quantitatively the strengths of the proposed model on several open-vocabulary tasks: Zero-shot 3D semantic segmentation using existing datasets; 3D grounding and retrieval of free-form language queries, using a small dataset that we propose as an extension of nuScenes. You can find the project page here https://vobecant.github.io/POP3D.
翻译:我们提出一种方法,从输入二维图像预测开词汇三维语义体素占据图,旨在实现自由形式语言查询的三维锚定、分割与检索。由于二维-三维歧义性以及目标任务的开词汇特性(获取三维标注训练数据十分困难),这一问题极具挑战性。本文做出三项贡献:第一,设计了一种用于开词汇三维语义占据预测的新型模型架构,该架构由二维-三维编码器结合占据预测与三维语言头构成,输出三维锚定语言嵌入的密集体素图,支持多种开词汇任务。第二,开发了一种三模态自监督学习算法,利用图像、语言与激光雷达点云三种模态,借助强预训练视觉语言模型训练所提架构,无需任何三维人工语言标注。第三,在多项开词汇任务上定量验证模型优势:利用现有数据集进行零样本三维语义分割;基于我们提出的nuScenes扩展小数据集,实现自由形式语言查询的三维锚定与检索。项目页面详见https://vobecant.github.io/POP3D。