We introduce Gaga, a framework that reconstructs and segments open-world 3D scenes by leveraging inconsistent 2D masks predicted by zero-shot segmentation models. Contrasted to prior 3D scene segmentation approaches that heavily rely on video object tracking, Gaga utilizes spatial information and effectively associates object masks across diverse camera poses. By eliminating the assumption of continuous view changes in training images, Gaga demonstrates robustness to variations in camera poses, particularly beneficial for sparsely sampled images, ensuring precise mask label consistency. Furthermore, Gaga accommodates 2D segmentation masks from diverse sources and demonstrates robust performance with different open-world zero-shot segmentation models, enhancing its versatility. Extensive qualitative and quantitative evaluations demonstrate that Gaga performs favorably against state-of-the-art methods, emphasizing its potential for real-world applications such as scene understanding and manipulation.
翻译:我们提出了Gaga框架,通过利用零样本分割模型预测的不一致二维掩码,实现对开放世界三维场景的重建与分割。与先前严重依赖视频目标追踪的三维场景分割方法不同,Gaga利用空间信息有效关联不同相机视角下的目标掩码。通过消除训练图像中连续视角变化的假设,Gaga展现出对相机视角变化的鲁棒性,尤其在稀疏采样图像中能确保掩码标签的一致性。此外,Gaga可兼容来自不同来源的二维分割掩码,并能与多种开放世界零样本分割模型稳定配合,显著增强了其通用性。大量定性与定量评估表明,Gaga的性能优于现有最先进方法,凸显其在场景理解与操作等实际应用中的潜力。