Diffusion-based video editing have reached impressive quality and can transform either the global style, local structure, and attributes of given video inputs, following textual edit prompts. However, such solutions typically incur heavy memory and computational costs to generate temporally-coherent frames, either in the form of diffusion inversion and/or cross-frame attention. In this paper, we conduct an analysis of such inefficiencies, and suggest simple yet effective modifications that allow significant speed-ups whilst maintaining quality. Moreover, we introduce Object-Centric Diffusion, coined as OCD, to further reduce latency by allocating computations more towards foreground edited regions that are arguably more important for perceptual quality. We achieve this by two novel proposals: i) Object-Centric Sampling, decoupling the diffusion steps spent on salient regions or background, allocating most of the model capacity to the former, and ii) Object-Centric 3D Token Merging, which reduces cost of cross-frame attention by fusing redundant tokens in unimportant background regions. Both techniques are readily applicable to a given video editing model \textit{without} retraining, and can drastically reduce its memory and computational cost. We evaluate our proposals on inversion-based and control-signal-based editing pipelines, and show a latency reduction up to 10x for a comparable synthesis quality.
翻译:基于扩散模型的视频编辑技术已取得令人瞩目的质量提升,能够根据文本编辑提示,对给定的视频输入进行全局风格、局部结构及属性的转换。然而,这类解决方案通常需要巨大的内存和计算成本来生成时间连贯的帧,具体表现为扩散反转和/或跨帧注意力机制的形式。本文对这类低效问题进行了深入分析,并提出简单而有效的改进方法,在保持编辑质量的同时实现显著的加速。此外,我们引入了一种名为“面向对象的扩散”(OCD)的方法,通过将计算资源更多地分配给对感知质量更重要的前景编辑区域,进一步降低延迟。为此我们提出两项创新方案:(i) 面向对象的采样,将扩散步骤在显著区域与背景区域间解耦,将大部分模型计算能力分配给前者;(ii) 面向对象的3D令牌合并,通过融合不重要的背景区域中的冗余令牌来降低跨帧注意力的计算成本。这两种技术均可直接应用于现有视频编辑模型而无需重新训练,并显著降低其内存与计算开销。我们在基于反转和基于控制信号的编辑流程上评估了所提方案,结果表明:在合成质量相当的情况下,延迟可降低达10倍。