Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene Transformer (DyST) model leverages recent work in neural scene representation to learn a latent decomposition of monocular real-world videos into scene content, per-view scene dynamics, and camera pose. This separation is achieved through a novel co-training scheme on monocular videos and our new synthetic dataset DySO. DyST learns tangible latent representations for dynamic scenes that enable view generation with separate control over the camera and the content of the scene.
翻译:对世界的视觉理解超越了单张图像的语义和平面结构。在本研究中,我们旨在从单目真实世界视频中捕捉场景的三维结构与动态信息。我们的动态场景Transformer(DyST)模型利用神经场景表示领域的最新成果,将单目真实世界视频的潜在分解学习为场景内容、逐视角场景动态以及相机位姿。这一分离通过一种新颖的联合训练方案实现,该方案基于单目视频与我们新合成的数据集DySO。DyST为动态场景学到了可解释的潜在表征,从而能够生成带有独立控制相机与场景内容的视角。