Face animation has achieved much progress in computer vision. However, prevailing GAN-based methods suffer from unnatural distortions and artifacts due to sophisticated motion deformation. In this paper, we propose a Face Animation framework with an attribute-guided Diffusion Model (FADM), which is the first work to exploit the superior modeling capacity of diffusion models for photo-realistic talking-head generation. To mitigate the uncontrollable synthesis effect of the diffusion model, we design an Attribute-Guided Conditioning Network (AGCN) to adaptively combine the coarse animation features and 3D face reconstruction results, which can incorporate appearance and motion conditions into the diffusion process. These specific designs help FADM rectify unnatural artifacts and distortions, and also enrich high-fidelity facial details through iterative diffusion refinements with accurate animation attributes. FADM can flexibly and effectively improve existing animation videos. Extensive experiments on widely used talking-head benchmarks validate the effectiveness of FADM over prior arts.
翻译:人脸动画在计算机视觉领域取得了长足进展。然而,基于生成对抗网络的主流方法因复杂的运动变形而存在不自然的扭曲和伪影。本文提出一种面向属性引导扩散模型的人脸动画框架(FADM),这是首个利用扩散模型卓越建模能力实现照片级逼真说话头部生成的工作。为缓解扩散模型不可控的合成效果,我们设计了属性引导条件网络(AGCN),通过自适应融合粗粒度动画特征与三维人脸重建结果,将外观和运动条件融入扩散过程。这些针对性设计不仅能帮助FADM修正不自然的伪影和扭曲,还能通过迭代式扩散精化与精准动画属性,丰富高保真面部细节。FADM可灵活高效地改善现有动画视频的质量。在广泛使用的说话头部基准数据集上的大量实验验证了FADM相较于现有方法的有效性。