We introduce Text2Immersion, an elegant method for producing high-quality 3D immersive scenes from text prompts. Our proposed pipeline initiates by progressively generating a Gaussian cloud using pre-trained 2D diffusion and depth estimation models. This is followed by a refining stage on the Gaussian cloud, interpolating and refining it to enhance the details of the generated scene. Distinct from prevalent methods that focus on single object or indoor scenes, or employ zoom-out trajectories, our approach generates diverse scenes with various objects, even extending to the creation of imaginary scenes. Consequently, Text2Immersion can have wide-ranging implications for various applications such as virtual reality, game development, and automated content creation. Extensive evaluations demonstrate that our system surpasses other methods in rendering quality and diversity, further progressing towards text-driven 3D scene generation. We will make the source code publicly accessible at the project page.
翻译:我们提出Text2Immersion,一种从文本提示生成高质量3D沉浸式场景的精妙方法。我们的流程首先利用预训练的2D扩散模型和深度估计模型逐步生成高斯点云,随后对高斯点云进行插值与细化处理以增强生成场景的细节。与聚焦于单一物体、室内场景或采用推拉镜头轨迹的现有方法不同,我们的方法能生成包含多样物体的各类场景,甚至可延伸至想象场景的创作。因此,Text2Immersion在虚拟现实、游戏开发及自动化内容生成等诸多应用中具有广泛前景。大量评估表明,我们的系统在渲染质量与多样性上超越其他方法,进一步推动了文本驱动的3D场景生成技术。源代码将在项目页面公开。