The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuity. Additionally, the efficiency of generating videos in pixel space is quite low.In this paper, we propose a novel approach to address these challenges by disentangling the target RGB pixels into two distinct components: spatial content and temporal motions. Specifically, we predict temporal motions which include motion vector and residual based on a 3D-UNet diffusion model. By explicitly modeling temporal motions and warping them to the starting image, we improve the temporal consistency of generated videos. This results in a reduction of spatial redundancy, emphasizing temporal details. Our proposed method achieves performance improvements by disentangling content and motion, all without introducing new structural complexities to the model. Extensive experiments on various datasets confirm our approach's superior performance over the majority of state-of-the-art methods in both effectiveness and efficiency.
翻译:条件式图像到视频生成(cI2V)的目标是通过给定条件(即一张图像和一段文本)创建可信的新视频。以往的cI2V生成方法通常在RGB像素空间中执行,在建模运动一致性和视觉连续性方面存在局限性。此外,在像素空间中生成视频的效率相当低下。本文提出了一种新颖方法应对这些挑战,通过将目标RGB像素解耦为两个独立分量:空间内容与时间运动。具体而言,我们基于3D-UNet扩散模型预测包含运动矢量和残差的时间运动。通过显式建模时间运动并将其扭曲至起始图像,我们提升了生成视频的时间一致性。这既减少了空间冗余,又突出了时间细节。所提方法在无需引入模型结构复杂性的前提下,通过内容与运动的解耦实现了性能提升。在多个数据集上的大量实验证实,本方法在有效性和效率上均优于大多数现有最优方法。