Instance segmentation in 3D is a challenging task due to the lack of large-scale annotated datasets. In this paper, we show that this task can be addressed effectively by leveraging instead 2D pre-trained models for instance segmentation. We propose a novel approach to lift 2D segments to 3D and fuse them by means of a neural field representation, which encourages multi-view consistency across frames. The core of our approach is a slow-fast clustering objective function, which is scalable and well-suited for scenes with a large number of objects. Unlike previous approaches, our method does not require an upper bound on the number of objects or object tracking across frames. To demonstrate the scalability of the slow-fast clustering, we create a new semi-realistic dataset called the Messy Rooms dataset, which features scenes with up to 500 objects per scene. Our approach outperforms the state-of-the-art on challenging scenes from the ScanNet, Hypersim, and Replica datasets, as well as on our newly created Messy Rooms dataset, demonstrating the effectiveness and scalability of our slow-fast clustering method.
翻译:三维实例分割因缺乏大规模标注数据集而极具挑战。本文表明,通过利用二维预训练实例分割模型可有效解决该任务。我们提出一种创新方法,将二维分割结果提升至三维空间,并通过神经场表征进行融合,从而增强跨帧的多视角一致性。该方法的核心是慢-快聚类目标函数,该函数具有可扩展性,特别适用于包含大量物体的场景。与以往方法不同,我们的方法无需设定物体数量上限或进行跨帧物体追踪。为验证慢-快聚类的可扩展性,我们构建了名为"凌乱房间"(Messy Rooms)的半真实数据集,其场景中每帧包含多达500个物体。我们的方法在ScanNet、Hypersim、Replica等数据集的复杂场景,以及新构建的Messy Rooms数据集上均超越现有最优方法,充分证明了慢-快聚类方法的有效性与可扩展性。