We propose $\texttt{SAL}$ ($\texttt{S}$egment $\texttt{A}$nything in $\texttt{L}$idar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training without manual supervision. While the established paradigm for $\textit{Lidar Panoptic Segmentation}$ (LPS) relies on manual supervision for a handful of object classes defined a priori, we utilize 2D vision foundation models to generate 3D supervision "for free". Our pseudo-labels consist of instance masks and corresponding CLIP tokens, which we lift to Lidar using calibrated multi-modal data. By training our model on these labels, we distill the 2D foundation models into our Lidar $\texttt{SAL}$ model. Even without manual labels, our model achieves $91\%$ in terms of class-agnostic segmentation and $44\%$ in terms of zero-shot LPS of the fully supervised state-of-the-art. Furthermore, we outperform several baselines that do not distill but only lift image features to 3D. More importantly, we demonstrate that $\texttt{SAL}$ supports arbitrary class prompts, can be easily extended to new datasets, and shows significant potential to improve with increasing amounts of self-labeled data.
翻译:我们提出$\texttt{SAL}$($\texttt{S}$egment $\texttt{A}$nything in $\texttt{L}$idar)方法,包含一个基于文本提示的零样本模型,用于在激光雷达中分割和分类任意目标,以及一个伪标签生成引擎,可在无人工监督下促进模型训练。尽管现有的$\textit{激光雷达全景分割}$(LPS)范式依赖于对预先定义的少数目标类别进行人工标注,但我们利用2D视觉基础模型“免费”生成3D监督信号。我们的伪标签由实例掩码及相应的CLIP标记组成,通过标定的多模态数据将其提升至激光雷达空间。通过在这些标注上训练模型,我们将2D基础模型的知识蒸馏到激光雷达$\texttt{SAL}$模型中。即使没有人工标签,我们的模型在类别无关分割任务上达到全监督最优方法的$91\%$,在零样本LPS任务上达到$44\%$。此外,我们优于多个仅将图像特征提升至3D而未进行知识蒸馏的基线方法。更重要的是,我们证明$\texttt{SAL}$支持任意类别提示,可轻松扩展至新数据集,并展现出随自标注数据量增加而显著提升的潜力。