Recently, talking face generation has drawn ever-increasing attention from the research community in computer vision due to its arduous challenges and widespread application scenarios, e.g. movie animation and virtual anchor. Although persevering efforts have been undertaken to enhance the fidelity and lip-sync quality of generated talking face videos, there is still large room for further improvements of synthesis quality and efficiency. Actually, these attempts somewhat ignore the explorations of fine-granularity feature extraction/integration and the consistency between probability distributions of landmarks, thereby recurring the issues of local details blurring and degraded fidelity. To mitigate these dilemmas, in this paper, a novel CLIP-based Attention and Probability Map Guided Network (CPNet) is delicately designed for inferring high-fidelity talking face videos. Specifically, considering the demands of fine-grained feature recalibration, a clip-based attention condenser is exploited to transfer knowledge with rich semantic priors from the prevailing CLIP model. Moreover, to guarantee the consistency in probability space and suppress the landmark ambiguity, we creatively propose the density map of facial landmark as auxiliary supervisory signal to guide the landmark distribution learning of generated frame. Extensive experiments on the widely-used benchmark dataset demonstrate the superiority of our CPNet against state of the arts in terms of image and lip-sync quality. In addition, a cohort of studies are also conducted to ablate the impacts of the individual pivotal components.
翻译:近年来,说话人脸生成因其严峻的挑战和广泛的应用场景(如电影动画和虚拟主播)而受到计算机视觉研究领域越来越多的关注。尽管为提高生成说话人脸视频的保真度和唇形同步质量付出了不懈努力,但合成质量与效率仍有较大提升空间。实际上,这些尝试在一定程度上忽视了细粒度特征提取/整合的探索以及地标概率分布一致性,从而导致局部细节模糊和保真度下降的问题。为缓解这些难题,本文精心设计了一种新颖的基于CLIP的注意力与概率图引导网络(CPNet),用于推断高保真说话人脸视频。具体而言,考虑到细粒度特征重校准的需求,我们利用基于CLIP的注意力凝聚器从主流的CLIP模型中转移富含语义先验的知识。此外,为保障概率空间的一致性并抑制地标歧义,我们创新性地提出人脸地标密度图作为辅助监督信号,引导生成帧的地标分布学习。在广泛使用的基准数据集上的大量实验表明,我们的CPNet在图像质量和唇形同步方面均优于现有最先进方法。同时,我们还进行了一系列研究以消融分析各个关键组成部分的影响。