Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

Emotional Talking Face synthesis is pivotal in multimedia and signal processing, yet existing 3D methods suffer from two critical challenges: poor audio-vision emotion alignment, manifested as difficult audio emotion extraction and inadequate control over emotional micro-expressions; and a one-size-fits-all multi-view fusion strategy that overlooks uncertainty and feature quality differences, undermining rendering quality. We propose UA-3DTalk, Uncertainty-Aware 3D Emotional Talking Face Synthesis with emotion prior distillation, which has three core modules: the Prior Extraction module disentangles audio into content-synchronized features for alignment and person-specific complementary features for individualization; the Emotion Distillation module introduces a multi-modal attention-weighted fusion mechanism and 4D Gaussian encoding with multi-resolution code-books, enabling fine-grained audio emotion extraction and precise control of emotional micro-expressions; the Uncertainty-based Deformation deploys uncertainty blocks to estimate view-specific aleatoric (input noise) and epistemic (model parameters) uncertainty, realizing adaptive multi-view fusion and incorporating a multi-head decoder for Gaussian primitive optimization to mitigate the limitations of uniform-weight fusion. Extensive experiments on regular and emotional datasets show UA-3DTalk outperforms state-of-the-art methods like DEGSTalk and EDTalk by 5.2% in E-FID for emotion alignment, 3.1% in SyncC for lip synchronization, and 0.015 in LPIPS for rendering quality. Project page: https://mrask999.github.io/UA-3DTalk

翻译：情感说话人脸合成在多媒体与信号处理领域至关重要，然而现有的三维方法面临两大关键挑战：一是音视频情感对齐效果不佳，表现为音频情感提取困难以及对情感微表情的控制不足；二是采用“一刀切”的多视角融合策略，忽视了不确定性和特征质量差异，从而损害了渲染质量。我们提出了UA-3DTalk，一种基于不确定性感知和情感先验蒸馏的三维情感说话人脸合成方法，其包含三个核心模块：先验提取模块将音频解耦为用于对齐的内容同步特征和用于个性化的个体互补特征；情感蒸馏模块引入了多模态注意力加权融合机制以及结合多分辨率码本的四维高斯编码，实现了细粒度的音频情感提取和对情感微表情的精确控制；基于不确定性的形变模块部署了不确定性块来估计视角特定的偶然性（输入噪声）和认知性（模型参数）不确定性，实现了自适应的多视角融合，并采用多头解码器对高斯基元进行优化，以缓解均匀权重融合的局限性。在常规数据集和情感数据集上进行的大量实验表明，UA-3DTalk在情感对齐的E-FID指标上优于DEGSTalk和EDTalk等最先进方法5.2%，在唇形同步的SyncC指标上提升3.1%，在渲染质量的LPIPS指标上提升0.015。项目页面：https://mrask999.github.io/UA-3DTalk