We introduce Open3DIS, a novel solution designed to tackle the problem of Open-Vocabulary Instance Segmentation within 3D scenes. Objects within 3D environments exhibit diverse shapes, scales, and colors, making precise instance-level identification a challenging task. Recent advancements in Open-Vocabulary scene understanding have made significant strides in this area by employing class-agnostic 3D instance proposal networks for object localization and learning queryable features for each 3D mask. While these methods produce high-quality instance proposals, they struggle with identifying small-scale and geometrically ambiguous objects. The key idea of our method is a new module that aggregates 2D instance masks across frames and maps them to geometrically coherent point cloud regions as high-quality object proposals addressing the above limitations. These are then combined with 3D class-agnostic instance proposals to include a wide range of objects in the real world. To validate our approach, we conducted experiments on three prominent datasets, including ScanNet200, S3DIS, and Replica, demonstrating significant performance gains in segmenting objects with diverse categories over the state-of-the-art approaches.
翻译:我们提出了Open3DIS,一种面向3D场景中开放词汇实例分割问题的新型解决方案。3D环境中的物体呈现多样的形状、尺度和颜色,使得精确的实例级识别成为一项具有挑战性的任务。近年来,开放词汇场景理解领域通过采用类别无关的3D实例提议网络进行目标定位,并为每个3D掩码学习可查询特征,取得了显著进展。尽管这些方法能生成高质量的实例提议,但在识别小尺度及几何模糊的物体时仍存在困难。本方法的核心创新在于设计了一个新模块,该模块跨帧聚合2D实例掩码,并将其映射为几何一致的3D点云区域,作为高质量物体提议以克服上述局限。随后将这些提议与3D类别无关的实例提议相结合,以覆盖真实世界中的广泛物体。为验证本方法,我们在ScanNet200、S3DIS和Replica三个主流数据集上开展了实验,结果表明在分割多样类别物体方面,本方法相比现有最先进技术取得了显著的性能提升。