Talking Face Generation (TFG) aims to reconstruct facial movements to achieve high natural lip movements from audio and facial features that are under potential connections. Existing TFG methods have made significant advancements to produce natural and realistic images. However, most work rarely takes visual quality into consideration. It is challenging to ensure lip synchronization while avoiding visual quality degradation in cross-modal generation methods. To address this issue, we propose a universal High-Definition Teeth Restoration Network, dubbed HDTR-Net, for arbitrary TFG methods. HDTR-Net can enhance teeth regions at an extremely fast speed while maintaining synchronization, and temporal consistency. In particular, we propose a Fine-Grained Feature Fusion (FGFF) module to effectively capture fine texture feature information around teeth and surrounding regions, and use these features to fine-grain the feature map to enhance the clarity of teeth. Extensive experiments show that our method can be adapted to arbitrary TFG methods without suffering from lip synchronization and frame coherence. Another advantage of HDTR-Net is its real-time generation ability. Also under the condition of high-definition restoration of talking face video synthesis, its inference speed is $300\%$ faster than the current state-of-the-art face restoration based on super-resolution.
翻译:说话人脸生成旨在从音频与面部特征中重建面部运动,以实现高度自然的口唇运动。现有TFG方法在生成自然逼真图像方面取得显著进展,但多数工作较少考虑视觉质量。跨模态生成方法在确保口唇同步的同时避免视觉质量退化仍具挑战性。为解决该问题,我们提出通用高清牙齿修复网络HDTR-Net,可适配任意TFG方法。HDTR-Net能在保持同步性与时序一致性的前提下,以极快速度增强牙齿区域。具体而言,我们提出细粒度特征融合模块,有效捕获牙齿及周围区域的精细纹理特征信息,并利用这些特征对特征图进行细粒度优化以提升牙齿清晰度。大量实验表明,本方法可适配任意TFG方法且不影响口唇同步与帧连贯性。HDTR-Net的另一优势在于其实时生成能力,在说话人脸视频合成的高清修复条件下,其推理速度比当前基于超分辨率的最先进人脸修复方法快300%。