We introduce animated stickers, a video diffusion model which generates an animation conditioned on a text prompt and static sticker image. Our model is built on top of the state-of-the-art Emu text-to-image model, with the addition of temporal layers to model motion. Due to the domain gap, i.e. differences in visual and motion style, a model which performed well on generating natural videos can no longer generate vivid videos when applied to stickers. To bridge this gap, we employ a two-stage finetuning pipeline: first with weakly in-domain data, followed by human-in-the-loop (HITL) strategy which we term ensemble-of-teachers. It distills the best qualities of multiple teachers into a smaller student model. We show that this strategy allows us to specifically target improvements to motion quality while maintaining the style from the static image. With inference optimizations, our model is able to generate an eight-frame video with high-quality, interesting, and relevant motion in under one second.
翻译:我们提出动画贴纸,一种基于文本提示和静态贴纸图像生成动画的视频扩散模型。该模型构建于最先进的Emu文本到图像模型之上,通过添加时间层来建模运动。由于领域差异(即视觉和运动风格的不同),原本在自然视频生成上表现良好的模型应用于贴纸时,无法再生成生动的视频。为弥合这一差距,我们采用两阶段微调流程:首先使用弱域内数据,随后采用我们称之为“教师集成”的人机协同(HITL)策略。该策略将多个教师模型的最佳特性蒸馏到较小的学生模型中。研究表明,此方法能够有针对性地提升运动质量,同时保持静态图像的风格。通过推理优化,我们的模型可在不到一秒内生成高质量、有趣且运动相关的八帧视频。