Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in inconsistencies across different parts of the face. For instance, Deepfake techniques may change smiling lips to an upset lip, while the eyes remain smiling. Existing detection methods depend on specific indicators of forgery, which tend to disappear as the forgery patterns are improved. To address the limitation, we propose Mover, a new Deepfake detection model that exploits unspecific facial part inconsistencies, which are inevitable weaknesses of Deepfake videos. Mover randomly masks regions of interest (ROIs) and recovers faces to learn unspecific features, which makes it difficult for fake faces to be recovered, while real faces can be easily recovered. Specifically, given a real face image, we first pretrain a masked autoencoder to learn facial part consistency by dividing faces into three parts and randomly masking ROIs, which are then recovered based on the unmasked facial parts. Furthermore, to maximize the discrepancy between real and fake videos, we propose a novel model with dual networks that utilize the pretrained encoder and masked autoencoder, respectively. 1) The pretrained encoder is finetuned for capturing the encoding of inconsistent information in the given video. 2) The pretrained masked autoencoder is utilized for mapping faces and distinguishing real and fake videos. Our extensive experiments on standard benchmarks demonstrate that Mover is highly effective.
翻译:深度伪造技术已被广泛用于恶意目的,促使研究者积极开发深度伪造检测方法。深度伪造操作通常涉及篡改面部局部区域,可能导致面部不同区域之间出现不一致性。例如,深度伪造技术可能将微笑的嘴唇改为不悦的嘴唇,而眼睛仍保持微笑状态。现有检测方法依赖于特定的伪造痕迹指标,但这些指标会随着伪造模式的改进而消失。为解决这一局限,我们提出Mover——一种利用非特定面部局部不一致性的深度伪造检测模型,这种不一致性是深度伪造视频不可避免的缺陷。Mover通过随机掩码感兴趣区域(ROI)并恢复人脸来学习非特定特征,使得伪造人脸难以被恢复,而真实人脸则能被轻松恢复。具体而言,给定真实人脸图像,我们首先预训练一个掩码自编码器,通过将人脸划分为三个区域并随机掩码ROI,再基于未掩码的面部区域进行恢复,从而学习面部局部一致性。此外,为最大化真实与伪造视频之间的差异,我们提出一种基于双网络的新颖模型,分别利用预训练编码器和掩码自编码器:1)微调预训练编码器以捕捉给定视频中不一致信息的编码;2)利用预训练掩码自编码器进行人脸映射及真实/伪造视频的判别。在标准基准数据集上的大量实验表明,Mover具有高效性。