The field of generative models has recently witnessed significant progress, with diffusion models showing remarkable performance in image generation. In light of this success, there is a growing interest in exploring the application of diffusion models to other modalities. One such challenge is the generation of coherent videos of complex scenes, which poses several technical difficulties, such as capturing temporal dependencies and generating long, high-resolution videos. This paper proposes GD-VDM, a novel diffusion model for video generation, demonstrating promising results. GD-VDM is based on a two-phase generation process involving generating depth videos followed by a novel diffusion Vid2Vid model that generates a coherent real-world video. We evaluated GD-VDM on the Cityscapes dataset and found that it generates more diverse and complex scenes compared to natural baselines, demonstrating the efficacy of our approach.
翻译:生成模型领域近期取得了显著进展,扩散模型在图像生成中展现出卓越性能。受此成功启发,研究人员开始探索扩散模型在其他模态中的应用。其中一项挑战是生成复杂场景的连贯视频,这涉及多个技术难点,例如捕捉时序依赖关系以及生成长时间、高分辨率的视频。本文提出GD-VDM,一种用于视频生成的新型扩散模型,并展示了其前景广阔的结果。GD-VDM基于两阶段生成过程:首先生成深度视频,随后通过一种新颖的扩散Vid2Vid模型生成连贯的真实世界视频。我们在Cityscapes数据集上对GD-VDM进行了评估,发现与自然基线方法相比,它能生成更多样化、更复杂的场景,证明了我们方法的有效性。