In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation that seldom puts emphasis on realistic head motions, we devise a GAN-based architecture that learns to synthesize rich head motion sequences over long duration while maintaining low error accumulation levels.In particular, the autoregressive generation of incremental outputs ensures smooth trajectories, while a multi-scale discriminator on input pairs drives generation toward better handling of high- and low-frequency signals and less mode collapse.We experimentally demonstrate the relevance of the proposed method and show its superiority compared to models that attained state-of-the-art performances on similar tasks.
翻译:在这项工作中,我们解决了从单一参考姿态出发,在低维语义空间中生成无条件头部运动以驱动静态人脸的课题。与通常不重视真实头部运动的传统音频条件驱动说话头生成不同,我们设计了一种基于GAN的架构,能够学习生成长时长的丰富头部运动序列,同时保持较低的误差累积水平。特别地,增量输出的自回归生成确保了平滑的轨迹,而输入对上的多尺度鉴别器则推动生成更优地处理高频和低频信号,并减少模式崩溃。我们通过实验证明了所提方法的相关性,并展示了其相较于在类似任务上达到最先进性能的模型的优越性。