This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remarkable ability to generate highly diverse videos following given text prompts, the fine-grained semantic edits are hard to control by plain textual prompt only in terms of object details and edited region, and cumbersome long text descriptions are usually needed for the task. We therefore investigate subject-driven video editing for more precise control of both edited regions and background preservation, and fine-grained semantic generation. We achieve this goal by introducing an reference image as supplementary input to the text-driven video editing, which avoids racking your brain to come up with a cumbersome text prompt describing the detailed appearance of the object. To limit the editing area, we refer to a method of cross attention control in image editing and successfully extend it to video editing by fusing the attention map of adjacent frames, which strikes a balance between maintaining video background and spatio-temporal consistency. Compared with current methods, the whole process of our method is like ``cut" the source object to be edited and then ``paste" the target object provided by reference image. We demonstrate that our method performs favorably over prior arts for video editing under the guidance of text prompt and extra reference image, as measured by both quantitative and subjective evaluations.
翻译:本文提出了一种名为Cut-and-Paste的新型框架,用于在文本提示和额外参考图像引导下的真实场景语义视频编辑。尽管文本驱动的视频编辑已展现出根据给定文本提示生成高度多样化视频的显著能力,但仅凭纯文本提示难以精细控制对象细节与编辑区域等细粒度语义编辑,且通常需要冗长的文本描述来完成该任务。为此,我们研究了主体驱动视频编辑,以实现对编辑区域和背景保留的更精确控制,以及细粒度语义生成。我们通过引入参考图像作为文本驱动视频编辑的补充输入来实现这一目标,从而避免绞尽脑汁编写描述对象详细外观的冗长文本提示。为限制编辑区域,我们借鉴了图像编辑中的交叉注意力控制方法,并通过融合相邻帧的注意力图成功将其扩展至视频编辑,在保持视频背景与时空一致性之间取得了平衡。与现有方法相比,我们的方法整个流程如同"剪切"待编辑的源对象,然后"粘贴"参考图像提供的目标对象。通过定量与主观评估,我们证明了该方法在文本提示与额外参考图像引导下的视频编辑任务中优于现有技术。