It is still an interesting and challenging problem to synthesize a vivid and realistic singing face driven by music signal. In this paper, we present a method for this task with natural motions of the lip, facial expression, head pose, and eye states. Due to the coupling of the mixed information of human voice and background music in common signals of music audio, we design a decouple-and-fuse strategy to tackle the challenge. We first decompose the input music audio into human voice stream and background music stream. Due to the implicit and complicated correlation between the two-stream input signals and the dynamics of the facial expressions, head motions and eye states, we model their relationship with an attention scheme, where the effects of the two streams are fused seamlessly. Furthermore, to improve the expressiveness of the generated results, we propose to decompose head movements generation into speed generation and direction generation, and decompose eye states generation into the short-time eye blinking generation and the long-time eye closing generation to model them separately. We also build a novel SingingFace Dataset to support the training and evaluation of this task, and to facilitate future works on this topic. Extensive experiments and user study show that our proposed method is capable of synthesizing vivid singing face, which is better than state-of-the-art methods qualitatively and quantitatively.
翻译:用音乐信号合成生动逼真的歌唱面部仍是一个有趣且富有挑战性的问题。本文提出一种方法,能够生成包含嘴唇、面部表情、头部姿态和眼睛状态自然运动的歌唱面部。针对常见音乐音频中人声与背景音乐混合信息相互耦合的问题,我们设计了一种解耦-融合策略来应对这一挑战。首先将输入音乐音频分解为人声流和背景音乐流。由于双流输入信号与面部表情、头部运动和眼睛状态动力学之间存在隐式且复杂的关联,我们采用注意力机制对其关系进行建模,将两路信号的影响无缝融合。此外,为提升生成结果的表现力,我们提出将头部运动生成分解为速度生成和方向生成,将眼睛状态生成分解为短时眨眼生成和长时闭眼生成分别建模。我们构建了新型SingingFace数据集,用于支持该任务的训练与评估,并为该领域的后续研究提供便利。大量实验和用户研究表明,本文提出的方法能够合成生动的歌唱面部,在定性和定量比较中均优于现有最优方法。