Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress. However, much of the effort has been put into lip syncing and rendering quality while the generation of natural head motion, let alone the audio-visual correlation between head motion and speech, has often been neglected. In this work, we propose a multi-scale audio-visual synchrony loss and a multi-scale autoregressive GAN to better handle short and long-term correlation between speech and the dynamics of the head and lips. In particular, we train a stack of syncer models on multimodal input pyramids and use these models as guidance in a multi-scale generator network to produce audio-aligned motion unfolding over diverse time scales. Our generator operates in the facial landmark domain, which is a standard low-dimensional head representation. The experiments show significant improvements over the state of the art in head motion dynamics quality and in multi-scale audio-visual synchrony both in the landmark domain and in the image domain.
翻译:使用深度生成模型通过语音输入信号对静态人脸图像进行动画化是一个活跃的研究课题,近年来取得了重要进展。然而,大部分努力集中于口型同步和渲染质量,而自然头部运动的生成——更不用说头部运动与语音之间的音视频相关性——往往被忽视。在这项工作中,我们提出了一种多尺度音视频同步损失函数和一种多尺度自回归生成对抗网络,以更好地处理语音与头部及嘴唇动态之间的短期和长期相关性。具体而言,我们在多模态输入金字塔上训练一组同步器模型,并将这些模型作为多尺度生成器网络中的引导机制,以产生在多样化时间尺度上与音频对齐的运动展开。我们的生成器在面部关键点域(一种标准的低维头部表示)中运行。实验表明,在关键点域和图像域中,我们的方法在头部运动动态质量以及多尺度音视频同步方面均显著优于现有技术。