The recent wave of AI-generated content has witnessed the great development and success of Text-to-Image (T2I) technologies. By contrast, Text-to-Video (T2V) still falls short of expectations though attracting increasing interests. Existing works either train from scratch or adapt large T2I model to videos, both of which are computation and resource expensive. In this work, we propose a Simple Diffusion Adapter (SimDA) that fine-tunes only 24M out of 1.1B parameters of a strong T2I model, adapting it to video generation in a parameter-efficient way. In particular, we turn the T2I model for T2V by designing light-weight spatial and temporal adapters for transfer learning. Besides, we change the original spatial attention to the proposed Latent-Shift Attention (LSA) for temporal consistency. With similar model architecture, we further train a video super-resolution model to generate high-definition (1024x1024) videos. In addition to T2V generation in the wild, SimDA could also be utilized in one-shot video editing with only 2 minutes tuning. Doing so, our method could minimize the training effort with extremely few tunable parameters for model adaptation.
翻译:近期AI生成内容的浪潮见证了文本到图像(T2I)技术的巨大发展与成功。相比之下,文本到视频(T2V)虽然吸引了越来越多的关注,但仍未达到预期。现有工作要么从头训练,要么将大型T2I模型适配到视频领域,两者都消耗大量计算和资源。本文提出了一种简单扩散适配器(SimDA),仅微调强T2I模型中11亿参数中的2400万,以参数高效的方式将其适配到视频生成。具体而言,我们通过设计轻量级空间和时间适配器进行迁移学习,将T2I模型转化为T2V模型。此外,我们将原始空间注意力改为所提出的潜移注意力(LSA)以确保时间一致性。采用相似的模型架构,我们还训练了一个视频超分辨率模型以生成高清(1024×1024)视频。除开放域T2V生成外,SimDA还可用于单镜头视频编辑,仅需2分钟微调。通过这种方式,我们的方法能以极少的可调参数最大程度降低训练成本,实现模型适配。