4D panoptic segmentation is a challenging but practically useful task that requires every point in a LiDAR point-cloud sequence to be assigned a semantic class label, and individual objects to be segmented and tracked over time. Existing approaches utilize only LiDAR inputs which convey limited information in regions with point sparsity. This problem can, however, be mitigated by utilizing RGB camera images which offer appearance-based information that can reinforce the geometry-based LiDAR features. Motivated by this, we propose 4D-Former: a novel method for 4D panoptic segmentation which leverages both LiDAR and image modalities, and predicts semantic masks as well as temporally consistent object masks for the input point-cloud sequence. We encode semantic classes and objects using a set of concise queries which absorb feature information from both data modalities. Additionally, we propose a learned mechanism to associate object tracks over time which reasons over both appearance and spatial location. We apply 4D-Former to the nuScenes and SemanticKITTI datasets where it achieves state-of-the-art results.
翻译:4D全景分割是一项具有挑战性但实用性强的任务,要求对激光雷达点云序列中的每个点赋予语义类别标签,并对个体目标进行时序分割与跟踪。现有方法仅利用激光雷达输入,在点稀疏区域信息有限。然而,通过引入RGB相机图像可缓解该问题,因为图像能提供基于外观的信息以增强基于几何的激光雷达特征。基于此,我们提出4D-Former:一种新颖的4D全景分割方法,同时利用激光雷达和图像模态,为输入点云序列预测语义掩码及时间一致的目标掩码。我们通过一组简洁的查询编码语义类别和目标,这些查询从两种数据模态中吸收特征信息。此外,我们提出一种基于学习的机制,通过结合外观与空间位置信息,对时序目标轨迹进行关联。我们在nuScenes和SemanticKITTI数据集上应用4D-Former,取得了最先进的结果。