Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has achieved rapid progress in recent years. However, most existing work focuses on generating lip movements only without handling the closely correlated facial expressions, which degrades the realism of the generated faces greatly. This paper presents DIRFA, a novel method that can generate talking faces with diverse yet realistic facial animations from the same driving audio. To accommodate fair variation of plausible facial animations for the same audio, we design a transformer-based probabilistic mapping network that can model the variational facial animation distribution conditioned upon the input audio and autoregressively convert the audio signals into a facial animation sequence. In addition, we introduce a temporally-biased mask into the mapping network, which allows to model the temporal dependency of facial animations and produce temporally smooth facial animation sequence. With the generated facial animation sequence and a source image, photo-realistic talking faces can be synthesized with a generic generation network. Extensive experiments show that DIRFA can generate talking faces with realistic facial animations effectively.
翻译:基于音频驱动的说话人脸生成旨在合成与音频对应的具有真实面部动画(包括精准的唇部运动、生动的面部表情细节和自然的头部姿态)的说话人脸,近年来取得了快速发展。然而,现有工作大多仅侧重于生成唇部运动,而未能处理与之密切相关的面部表情,这极大地降低了生成人脸的逼真度。本文提出了一种名为DIRFA的新方法,能够从同一驱动音频中生成具有多样且真实面部动画的说话人脸。为了适应同一音频下合理面部动画的公平变化,我们设计了一种基于Transformer的概率映射网络,该网络能够对输入音频条件下的变化性面部动画分布进行建模,并自回归地将音频信号转换为面部动画序列。此外,我们在映射网络中引入了一个时间偏置掩码,使得能够对面部动画的时间依赖性进行建模,并生成时间上平滑的面部动画序列。利用生成的面部动画序列和源图像,可通过通用生成网络合成照片级逼真的说话人脸。大量实验表明,DIRFA能够有效生成具有真实面部动画的说话人脸。