Unconditional video generation is a challenging task that involves synthesizing high-quality videos that are both coherent and of extended duration. To address this challenge, researchers have used pretrained StyleGAN image generators for high-quality frame synthesis and focused on motion generator design. The motion generator is trained in an autoregressive manner using heavy 3D convolutional discriminators to ensure motion coherence during video generation. In this paper, we introduce a novel motion generator design that uses a learning-based inversion network for GAN. The encoder in our method captures rich and smooth priors from encoding images to latents, and given the latent of an initially generated frame as guidance, our method can generate smooth future latent by modulating the inversion encoder temporally. Our method enjoys the advantage of sparse training and naturally constrains the generation space of our motion generator with the inversion network guided by the initial frame, eliminating the need for heavy discriminators. Moreover, our method supports style transfer with simple fine-tuning when the encoder is paired with a pretrained StyleGAN generator. Extensive experiments conducted on various benchmarks demonstrate the superiority of our method in generating long and high-resolution videos with decent single-frame quality and temporal consistency.
翻译:无条件视频生成是一项具有挑战性的任务,涉及合成高质量、连贯且时长延长的视频。为应对这一挑战,研究者借助预训练的StyleGAN图像生成器实现高质量帧合成,并聚焦于运动生成器设计。该运动生成器通过自回归方式训练,并采用重型三维卷积判别器确保视频生成过程中的运动连贯性。本文提出一种基于学习型GAN反演网络的新型运动生成器设计。该方法中的编码器通过将图像编码为潜在编码,捕获丰富且平滑的先验信息;在初始生成帧的潜在编码引导下,通过时序调制反演编码器,可生成平滑的未来潜在表示。本方法兼具稀疏训练的优势,并通过初始帧引导的反演网络自然约束运动生成器的生成空间,从而无需重型判别器。此外,当编码器与预训练StyleGAN生成器配对时,本方法可通过简单微调支持风格迁移。在多个基准测试上的大量实验表明,本方法在生成长时长、高分辨率视频方面具有优越性,且能保持帧内质量与时序一致性。