In this work, we present SceneDreamer, an unconditional generative model for unbounded 3D scenes, which synthesizes large-scale 3D landscapes from random noises. Our framework is learned from in-the-wild 2D image collections only, without any 3D annotations. At the core of SceneDreamer is a principled learning paradigm comprising 1) an efficient yet expressive 3D scene representation, 2) a generative scene parameterization, and 3) an effective renderer that can leverage the knowledge from 2D images. Our framework starts from an efficient bird's-eye-view (BEV) representation generated from simplex noise, which consists of a height field and a semantic field. The height field represents the surface elevation of 3D scenes, while the semantic field provides detailed scene semantics. This BEV scene representation enables 1) representing a 3D scene with quadratic complexity, 2) disentangled geometry and semantics, and 3) efficient training. Furthermore, we propose a novel generative neural hash grid to parameterize the latent space given 3D positions and the scene semantics, which aims to encode generalizable features across scenes. Lastly, a neural volumetric renderer, learned from 2D image collections through adversarial training, is employed to produce photorealistic images. Extensive experiments demonstrate the effectiveness of SceneDreamer and superiority over state-of-the-art methods in generating vivid yet diverse unbounded 3D worlds.
翻译:本文提出SceneDreamer——一种面向无界三维场景的无条件生成模型,该模型能够从随机噪声中合成大规模三维景观。我们的框架仅从自然场景的二维图像集合中学习,无需任何三维标注。SceneDreamer的核心是一个原则性学习范式,包含:1)高效且富有表现力的三维场景表示,2)生成式场景参数化,3)能够利用二维图像知识的有效渲染器。本框架始于由单纯形噪声生成的鸟瞰图(BEV)表示,该表示由高度场和语义场构成。高度场表征三维场景的地表高程,而语义场提供细粒度的场景语义信息。这种BEV场景表示具有以下优势:1)以二次复杂度表征三维场景,2)实现几何与语义的解耦,3)支持高效训练。此外,我们提出一种新型生成式神经哈希网格,用于基于三维位置和场景语义对潜在空间进行参数化,旨在编码跨场景的可泛化特征。最后,通过对抗训练从二维图像集合中学习的神经体积渲染器被用于生成逼真图像。大量实验证明了SceneDreamer的有效性,以及其在生成生动且多样化的无界三维世界方面相比现有最优方法的优越性。