We present "SemCity," a 3D diffusion model for semantic scene generation in real-world outdoor environments. Most 3D diffusion models focus on generating a single object, synthetic indoor scenes, or synthetic outdoor scenes, while the generation of real-world outdoor scenes is rarely addressed. In this paper, we concentrate on generating a real-outdoor scene through learning a diffusion model on a real-world outdoor dataset. In contrast to synthetic data, real-outdoor datasets often contain more empty spaces due to sensor limitations, causing challenges in learning real-outdoor distributions. To address this issue, we exploit a triplane representation as a proxy form of scene distributions to be learned by our diffusion model. Furthermore, we propose a triplane manipulation that integrates seamlessly with our triplane diffusion model. The manipulation improves our diffusion model's applicability in a variety of downstream tasks related to outdoor scene generation such as scene inpainting, scene outpainting, and semantic scene completion refinements. In experimental results, we demonstrate that our triplane diffusion model shows meaningful generation results compared with existing work in a real-outdoor dataset, SemanticKITTI. We also show our triplane manipulation facilitates seamlessly adding, removing, or modifying objects within a scene. Further, it also enables the expansion of scenes toward a city-level scale. Finally, we evaluate our method on semantic scene completion refinements where our diffusion model enhances predictions of semantic scene completion networks by learning scene distribution. Our code is available at https://github.com/zoomin-lee/SemCity.
翻译:摘要:我们提出“SemCity”,一种用于真实世界户外环境中语义场景生成的三维扩散模型。大多数三维扩散模型专注于生成单个物体、合成室内场景或合成室外场景,而真实世界户外场景的生成鲜有涉及。本文通过在实际户外数据集上学习扩散模型,集中探讨真实户外场景的生成。与合成数据不同,真实户外数据集因传感器限制通常包含更多空白区域,导致学习真实户外分布面临挑战。为解决此问题,我们利用三平面表示作为场景分布的代理形式,供扩散模型学习。此外,我们提出一种与三平面扩散模型无缝集成的三平面操作用于改进扩散模型在户外场景生成相关下游任务中的适用性,如场景修复、场景外延和语义场景补全优化。实验结果表明,在真实户外数据集SemanticKITTI上,我们的三平面扩散模型相较于现有工作生成了有意义的结果。我们还展示了三平面操作可实现场景内物体的无缝添加、移除或修改,并能将场景扩展至城市级规模。最后,我们在语义场景补全优化任务中评估了该方法:扩散模型通过学习场景分布增强了语义场景补全网络的预测效果。我们的代码已开源至https://github.com/zoomin-lee/SemCity。