We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during training. Existing methods are typically constrained to identifying simple objects, primarily due to insufficient object priors in the learning process. In this paper, we present FoundObj, a novel framework featuring a superpoint-based object discovery agent that incrementally merges suitable neighboring superpoints, guided by our innovative semantic and geometric reward modules. These modules synergistically leverage semantic and geometric priors from self-supervised 2D/3D foundation models, providing complementary feedback to the object discovery agent and enabling robust identification of multi-class objects through reinforcement learning. Extensive experiments on diverse benchmarks demonstrate that our approach consistently outperforms existing baselines. Notably, our method exhibits strong generalization in zero-shot and long-tail scenarios, underscoring its potential for scalable, label-free 3D object segmentation.
翻译:摘要: 本文针对复杂场景点云中无需依赖任何场景级人工标注训练的三维物体分割这一挑战性任务展开研究。现有方法通常局限于识别简单物体,根本原因在于学习过程中物体先验信息不足。我们提出FoundObj框架,该框架创新性地将超点作为基础,构建基于超点的物体发现智能体,通过逐步融合相邻的合适超点实现目标,并由所提出的语义与几何奖励模块协同引导。这些模块充分利用自监督2D/3D基础模型中的语义与几何先验,为物体发现智能体提供互补性反馈,借助强化学习实现多类物体的稳健识别。在多个基准数据集上的广泛实验表明,本方法持续优于现有基线方法。值得注意的是,本方法在零样本与长尾场景中展现出强大的泛化能力,充分彰显其在可扩展、无标注三维物体分割领域的巨大潜力。