Deepfake technologies empowered by deep learning are rapidly evolving, creating new security concerns for society. Existing multimodal detection methods usually capture audio-visual inconsistencies to expose Deepfake videos. More seriously, the advanced Deepfake technology realizes the audio-visual calibration of the critical phoneme-viseme regions, achieving a more realistic tampering effect, which brings new challenges. To address this problem, we propose a novel Deepfake detection method to mine the correlation between Non-critical Phonemes and Visemes, termed NPVForensics. Firstly, we propose the Local Feature Aggregation block with Swin Transformer (LFA-ST) to construct non-critical phoneme-viseme and corresponding facial feature streams effectively. Secondly, we design a loss function for the fine-grained motion of the talking face to measure the evolutionary consistency of non-critical phoneme-viseme. Next, we design a phoneme-viseme awareness module for cross-modal feature fusion and representation alignment, so that the modality gap can be reduced and the intrinsic complementarity of the two modalities can be better explored. Finally, a self-supervised pre-training strategy is leveraged to thoroughly learn the audio-visual correspondences in natural videos. In this manner, our model can be easily adapted to the downstream Deepfake datasets with fine-tuning. Extensive experiments on existing benchmarks demonstrate that the proposed approach outperforms state-of-the-art methods.
翻译:基于深度学习的深度伪造技术正在快速发展,给社会带来了新的安全隐患。现有跨模态检测方法通常通过捕捉音视频不一致性来揭露深度伪造视频。更严重的是,先进深度伪造技术实现了关键音素-视位区域的音视频校准,达到更逼真的篡改效果,这带来了新的挑战。针对此问题,我们提出一种新颖的深度伪造检测方法,挖掘非关键音素与视位之间的相关性,命名为NPVForensics。首先,我们提出基于Swin Transformer的局部特征聚合模块(LFA-ST),有效构建非关键音素-视位及其对应面部特征流。其次,我们设计面向说话人脸细粒度运动的损失函数,以衡量非关键音素-视位的演化一致性。接着,我们设计音素-视位感知模块用于跨模态特征融合与表征对齐,从而缩小模态差异并更充分探索两种模态的内在互补性。最后,利用自监督预训练策略充分学习自然视频中的音视频对应关系。通过这种方式,我们的模型可通过微调轻松迁移至下游深度伪造数据集。在现有基准上的大量实验表明,所提方法优于当前最优技术。