Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and quality. Extensive experiments show that Upscale-A-Video surpasses existing methods in both synthetic and real-world benchmarks, as well as in AI-generated videos, showcasing impressive visual realism and temporal consistency.
翻译:基于文本的扩散模型在生成和编辑任务中取得了显著成功,其生成先验为增强视觉内容展现了巨大潜力。然而,由于对输出保真度和时间一致性的高要求,且扩散模型固有的随机性增加了复杂性,将这些模型应用于视频超分辨率仍具挑战。本研究提出Upscale-A-Video——一种面向视频放大的文本引导潜在扩散框架。该框架通过两种关键机制确保时间一致性:在局部层面,将时间层集成到U-Net和VAE解码器中,维持短序列内的一致性;在全局层面,无需训练,引入流引导的循环潜在传播模块,通过跨整个序列传播和融合潜在表示来增强整体视频稳定性。得益于扩散范式,我们的模型还具备更高灵活性:允许文本提示引导纹理生成,并通过可调节噪声水平平衡恢复与生成,实现保真度与质量之间的权衡。大量实验表明,Upscale-A-Video在合成与真实基准测试以及AI生成视频任务中均优于现有方法,展现了出色的视觉真实感和时间一致性。