Latent Diffusion Models (LDMs) are renowned for their powerful capabilities in image and video synthesis. Yet, video editing methods suffer from insufficient pre-training data or video-by-video re-training cost. In addressing this gap, we propose FLDM (Fused Latent Diffusion Model), a training-free framework to achieve text-guided video editing by applying off-the-shelf image editing methods in video LDMs. Specifically, FLDM fuses latents from an image LDM and an video LDM during the denoising process. In this way, temporal consistency can be kept with video LDM while high-fidelity from the image LDM can also be exploited. Meanwhile, FLDM possesses high flexibility since both image LDM and video LDM can be replaced so advanced image editing methods such as InstructPix2Pix and ControlNet can be exploited. To the best of our knowledge, FLDM is the first method to adapt off-the-shelf image editing methods into video LDMs for video editing. Extensive quantitative and qualitative experiments demonstrate that FLDM can improve the textual alignment and temporal consistency of edited videos.
翻译:潜扩散模型(LDM)因其在图像和视频合成领域的强大能力而备受关注。然而,现有视频编辑方法面临预训练数据不足或逐视频重新训练成本高昂的问题。为弥补这一不足,本文提出FLDM(融合潜扩散模型),一种无需训练即可将现成图像编辑方法应用于视频LDM实现文本引导视频编辑的框架。具体而言,FLDM在去噪过程中融合来自图像LDM和视频LDM的潜变量,既能通过视频LDM保持时间一致性,又能利用图像LDM实现高保真度。同时,由于图像LDM和视频LDM均可替换,FLDM具备高度灵活性,可引入如InstructPix2Pix和ControlNet等先进图像编辑方法。据我们所知,FLDM是首个将现成图像编辑方法适配至视频LDM进行视频编辑的技术。大量定量与定性实验表明,FLDM能有效提升编辑视频的文本对齐程度与时间一致性。