We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such videos in an online fashion by combining the generative power of a pre-trained text-to-image model with the geometric priors learned by a pre-trained monocular depth prediction model. To tackle the pivotal challenge of achieving 3D consistency, i.e., synthesizing videos that depict geometrically-plausible scenes, we deploy an online test-time training to encourage the predicted depth map of the current frame to be geometrically consistent with the synthesized scene. The depth maps are used to construct a unified mesh representation of the scene, which is progressively constructed along the video generation process. In contrast to previous works, which are applicable only to limited domains, our method generates diverse scenes, such as walkthroughs in spaceships, caves, or ice castles.
翻译:我们提出了一种文本驱动的无限视角生成方法——仅根据描述场景的输入文本提示和相机位姿,即可合成各种场景的长时视频。我们引入了一个新颖的框架,通过结合预训练文本到图像模型的生成能力与预训练单目深度预测模型学到的几何先验,以在线方式生成此类视频。为解决实现三维一致性的关键挑战,即合成描绘几何合理场景的视频,我们采用了在线测试时训练,鼓励当前帧的预测深度图与合成场景在几何上保持一致。深度图用于构建场景的统一网格表示,该表示在视频生成过程中逐步构建。与仅适用于有限领域的先前工作不同,我们的方法能够生成多样化的场景,例如在太空飞船、洞穴或冰城堡中的漫游。