Most deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction, using a fully self-supervised approach. To this end, we design a trainable model that employs both incomplete 3D reconstructions and their corresponding source RGB-D images, fusing cross-domain features into volumetric embeddings to predict complete 3D geometry, color, and semantics. Our key technical innovation is to leverage differentiable rendering of color and semantics, using the observed RGB images and a generic semantic segmentation model as color and semantics supervision, respectively. We additionally develop a method to synthesize an augmented set of virtual training views complementing the original real captures, enabling more efficient self-supervision for semantics. In this work we propose an end-to-end trainable solution jointly addressing geometry completion, colorization, and semantic mapping from a few RGB-D images, without 3D or 2D ground-truth. Our method is the first, to our knowledge, fully self-supervised method addressing completion and semantic segmentation of real-world 3D scans. It performs comparably well with the 3D supervised baselines, surpasses baselines with 2D supervision on real datasets, and generalizes well to unseen scenes.
翻译:大多数对3D室内空间进行综合语义建模的深度学习方法,都需要在3D领域中进行昂贵的密集标注。本文探索了一种核心的3D场景建模任务——语义场景重建,采用完全自监督的方法实现。为此,我们设计了一个可训练模型,该模型同时利用不完整的3D重建结果及其对应的RGB-D源图像,将跨域特征融合为体积嵌入,以预测完整的3D几何、颜色和语义信息。我们的关键技术创新在于利用颜色和语义的可微分渲染,分别以观测到的RGB图像和通用语义分割模型作为颜色和语义的监督信号。此外,我们还开发了一种方法,用于合成一组增强的虚拟训练视角,以补充原始真实捕获数据,从而实现更高效的语义自监督。本文提出了一种端到端的可训练解决方案,可从少量RGB-D图像中联合解决几何补全、着色和语义映射问题,无需任何3D或2D真实标注。据我们所知,我们的方法是首个完全自监督的方法,能够处理真实世界3D扫描的补全和语义分割任务。该方法在性能上与基于3D监督的基线方法相当,在真实数据集上超越了基于2D监督的基线方法,并且对未见场景具有良好的泛化能力。