Recent advancements in diffusion models have unlocked unprecedented abilities in visual creation. However, current text-to-video generation models struggle with the trade-off among movement range, action coherence and object consistency. To mitigate this issue, we present a controllable text-to-video (T2V) diffusion model, called Control-A-Video, capable of maintaining consistency while customizable video synthesis. Based on a pre-trained conditional text-to-image (T2I) diffusion model, our model aims to generate videos conditioned on a sequence of control signals, such as edge or depth maps. For the purpose of improving object consistency, Control-A-Video integrates motion priors and content priors into video generation. We propose two motion-adaptive noise initialization strategies, which are based on pixel residual and optical flow, to introduce motion priors from input videos, producing more coherent videos. Moreover, a first-frame conditioned controller is proposed to generate videos from content priors of the first frame, which facilitates the semantic alignment with text and allows longer video generation in an auto-regressive manner. With the proposed architecture and strategies, our model achieves resource-efficient convergence and generate consistent and coherent videos with fine-grained control. Extensive experiments demonstrate its success in various video generative tasks such as video editing and video style transfer, outperforming previous methods in terms of consistency and quality.
翻译:扩散模型的最新进展在视觉创作领域解锁了前所未有的能力。然而,当前的文本到视频生成模型在运动范围、动作连贯性与对象一致性之间存在权衡困境。为解决这一问题,我们提出了一种名为Control-A-Video的可控文本到视频(T2V)扩散模型,该模型能够在保持一致性的同时实现可定制的视频合成。基于预训练的条件文本到图像(T2I)扩散模型,我们的模型旨在根据一系列控制信号(如边缘或深度图)生成视频。为提升对象一致性,Control-A-Video将运动先验与内容先验整合到视频生成中。我们提出了两种基于像素残差和光流的运动自适应噪声初始化策略,用于引入输入视频中的运动先验,从而生成更连贯的视频。此外,我们设计了一种首帧条件控制器,通过首帧的内容先验生成视频,这促进了与文本的语义对齐,并支持以自回归方式生成更长视频。借助所提出的架构与策略,我们的模型实现了资源高效的收敛,并生成了具有细粒度控制的一致且连贯的视频。大量实验证明了该模型在视频编辑、视频风格迁移等多种视频生成任务中的成功,在一致性与质量方面均优于先前方法。