Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving. DiffTalker addresses the challenges associated with directly applying diffusion models to audio control, which are traditionally trained on text-image pairs. DiffTalker consists of two agent networks: a transformer-based landmarks completion network for geometric accuracy and a diffusion-based face generation network for texture details. Landmarks play a pivotal role in establishing a seamless connection between the audio and image domains, facilitating the incorporation of knowledge from pre-trained diffusion models. This innovative approach efficiently produces articulate-speaking faces. Experimental results showcase DiffTalker's superior performance in producing clear and geometrically accurate talking faces, all without the need for additional alignment between audio and image features.
翻译:生成逼真的说话人脸是一项复杂且被广泛讨论的任务,具有众多应用。本文提出DiffTalker,一种通过音频和地标联合驱动生成栩栩如生说话人脸的新模型。DiffTalker解决了将扩散模型直接应用于音频控制时面临的挑战,这些模型传统上是在文本-图像对上训练的。DiffTalker由两个代理网络组成:基于Transformer的地标补全网络用于几何精度,以及基于扩散的人脸生成网络用于纹理细节。地标在音频和图像域之间建立无缝连接中发挥关键作用,有助于利用预训练扩散模型的知识。这种创新方法高效地生成表达清晰的说话人脸。实验结果表明,DiffTalker在生成清晰且几何精度高的说话人脸方面表现出色,且无需音频与图像特征之间的额外对齐。