3D scene generation has quickly become a challenging new research direction, fueled by consistent improvements of 2D generative diffusion models. Most prior work in this area generates scenes by iteratively stitching newly generated frames with existing geometry. These works often depend on pre-trained monocular depth estimators to lift the generated images into 3D, fusing them with the existing scene representation. These approaches are then often evaluated via a text metric, measuring the similarity between the generated images and a given text prompt. In this work, we make two fundamental contributions to the field of 3D scene generation. First, we note that lifting images to 3D with a monocular depth estimation model is suboptimal as it ignores the geometry of the existing scene. We thus introduce a novel depth completion model, trained via teacher distillation and self-training to learn the 3D fusion process, resulting in improved geometric coherence of the scene. Second, we introduce a new benchmarking scheme for scene generation methods that is based on ground truth geometry, and thus measures the quality of the structure of the scene.
翻译:三维场景生成已迅速成为一项富有挑战性的新研究方向,这得益于二维生成扩散模型的持续改进。该领域的大多数先前工作通过迭代地将新生成的帧与现有几何结构拼接来生成场景。这些方法通常依赖预训练的单目深度估计器将生成的图像提升至三维,并将其与现有场景表示融合。随后,这类方法常通过文本指标进行评估,即衡量生成图像与给定文本提示之间的相似度。本文对三维场景生成领域做出两项基础性贡献:首先,我们指出利用单目深度估计模型将图像提升至三维的做法并非最优,因为它忽略了现有场景的几何结构。为此,我们提出一种新型深度补全模型,通过教师蒸馏和自训练学习三维融合过程,从而提升场景的几何一致性。其次,我们引入了一种基于真实几何结构的场景生成方法新基准方案,以此衡量场景结构的质量。