Recent one-shot video tuning methods, which fine-tune the network on a specific video based on pre-trained text-to-image models (e.g., Stable Diffusion), are popular in the community because of the flexibility. However, these methods often produce videos marred by incoherence and inconsistency. To address these limitations, this paper introduces a simple yet effective noise constraint across video frames. This constraint aims to regulate noise predictions across their temporal neighbors, resulting in smooth latents. It can be simply included as a loss term during the training phase. By applying the loss to existing one-shot video tuning methods, we significantly improve the overall consistency and smoothness of the generated videos. Furthermore, we argue that current video evaluation metrics inadequately capture smoothness. To address this, we introduce a novel metric that considers detailed features and their temporal dynamics. Experimental results validate the effectiveness of our approach in producing smoother videos on various one-shot video tuning baselines. The source codes and video demos are available at \href{https://github.com/SPengLiang/SmoothVideo}{https://github.com/SPengLiang/SmoothVideo}.
翻译:近期基于单次视频的微调方法(即在预训练文本到图像模型(如Stable Diffusion)基础上对特定视频进行网络微调)因其灵活性广受学界关注。然而,此类方法生成的视频常存在时空不连贯与不一致性问题。针对上述局限,本文提出一种简洁高效的跨视频帧噪声约束机制,通过规范视频帧在时序邻域内的噪声预测行为,实现潜在表征的平滑过渡。该约束机制可简便地以损失函数形式嵌入训练阶段。将本损失函数应用于现有单次视频微调方法后,生成视频的整体一致性与平滑度得到显著提升。此外,本文指出现有视频评估指标难以充分表征平滑性,为此引入了一项新型评估指标,该指标综合考虑细节特征及其时序动态特性。实验结果表明,本方法在多种单次视频微调基线模型上均能有效提升视频平滑度。相关源代码与视频演示已公开于\href{https://github.com/SPengLiang/SmoothVideo}{https://github.com/SPengLiang/SmoothVideo}。