We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two stages with explicit motion modeling. For the first stage, we propose a diffusion-based motion field predictor, which focuses on deducing the trajectories of the reference image's pixels. For the second stage, we propose motion-augmented temporal attention to enhance the limited 1-D temporal attention in video latent diffusion models. This module can effectively propagate reference image's feature to synthesized frames with the guidance of predicted trajectories from the first stage. Compared with existing methods, Motion-I2V can generate more consistent videos even at the presence of large motion and viewpoint variation. By training a sparse trajectory ControlNet for the first stage, Motion-I2V can support users to precisely control motion trajectories and motion regions with sparse trajectory and region annotations. This offers more controllability of the I2V process than solely relying on textual instructions. Additionally, Motion-I2V's second stage naturally supports zero-shot video-to-video translation. Both qualitative and quantitative comparisons demonstrate the advantages of Motion-I2V over prior approaches in consistent and controllable image-to-video generation. Please see our project page at https://xiaoyushi97.github.io/Motion-I2V/.
翻译:我们提出了Motion-I2V,一种用于一致可控图像到视频生成的新型框架。与以往直接学习复杂图像到视频映射的方法不同,Motion-I2V通过显式运动建模将图像到视频生成分解为两个阶段。在第一阶段,我们提出了一种基于扩散模型的运动场预测器,专注于推断参考图像像素的运动轨迹。在第二阶段,我们提出了运动增强时间注意力机制,以增强视频潜伏扩散模型中有限的1D时间注意力。该模块能够在预测轨迹(来自第一阶段)引导下,将参考图像特征有效传播至合成帧。与现有方法相比,Motion-I2V即使在存在大运动和视角变化的情况下也能生成更一致的视频。通过为第一阶段训练稀疏轨迹ControlNet,Motion-I2V能够支持用户利用稀疏轨迹和区域标注精确控制运动轨迹与运动区域,这比仅依赖文本指令提供了更强的I2V过程可控性。此外,Motion-I2V的第二阶段天然支持零样本视频到视频转换。定性与定量比较均证明了Motion-I2V在一致可控图像到视频生成中相较于先前方法的优势。详见项目页面https://xiaoyushi97.github.io/Motion-I2V/。