With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial domain shift, substantially degrading detection performance. We construct the Singing Head DeepFake (SHDF) dataset using rhythm-aware generative models to fill the gap in singing benchmarks. To cope with cross-scenario domain shifts, we propose a Text-guided Audio-Visual Forgery Detection (T-AVFD) framework that generalizes across both talking and singing scenarios. T-AVFD comprises a facial authenticity pattern learner and a multi-modal differential weight learning module. The pattern learner aligns facial features with multi-granularity textual descriptions to learn generalizable authenticity patterns. The weight learning module preserves intrinsic audio-visual consistency and adaptively integrates it with authenticity patterns via differential weighting. Extensive experiments on multiple talking head deepfake datasets and SHDF show consistent improvements over existing baselines and strong robustness under diverse perturbations.
翻译:随着音视频生成模型的快速发展,可靠的伪造检测变得愈发关键。现有音视频深度伪造检测方法通常依赖跨模态不一致性。在唱歌场景中,节奏性发声削弱了这种耦合关系并引入了显著的域偏移,导致检测性能大幅下降。我们利用节奏感知生成模型构建了唱歌头部深度伪造(SHDF)数据集,以填补唱歌基准测试的空白。为应对跨场景域偏移,我们提出文本引导的音视频伪造检测(T-AVFD)框架,该框架可泛化至说话与唱歌两种场景。T-AVFD包含面部真实性模式学习器与多模态差异权重学习模块。模式学习器将面部特征与多粒度文本描述对齐,以学习可泛化的真实性模式;权重学习模块保留固有音视频一致性,并通过差异加权将其与真实性模式自适应融合。在多个说话头部深度伪造数据集及SHDF上的大量实验表明,该方法相较于现有基线取得持续改进,并在多种扰动下展现出强鲁棒性。