The emergence of deepfake technologies has become a matter of social concern as they pose threats to individual privacy and public security. It is now of great significance to develop reliable deepfake detectors. However, with numerous face manipulation algorithms present, it is almost impossible to collect sufficient representative fake faces, and it is hard for existing detectors to generalize to all types of manipulation. Therefore, we turn to learn the distribution of real faces, and indirectly identify fake images that deviate from the real face distribution. In this study, we propose Real Face Foundation Representation Learning (RFFR), which aims to learn a general representation from large-scale real face datasets and detect potential artifacts outside the distribution of RFFR. Specifically, we train a model on real face datasets by masked image modeling (MIM), which results in a discrepancy between input faces and the reconstructed ones when applying the model on fake samples. This discrepancy reveals the low-level artifacts not contained in RFFR, making it easier to build a deepfake detector sensitive to all kinds of potential artifacts outside the distribution of RFFR. Extensive experiments demonstrate that our method brings about better generalization performance, as it significantly outperforms the state-of-the-art methods in cross-manipulation evaluations, and has the potential to further improve by introducing extra real faces for training RFFR.
翻译:深度伪造技术的出现已成为社会关注的问题,因其对个人隐私和公共安全构成威胁。如今开发可靠的深度伪造检测器具有重要意义。然而,面对众多人脸篡改算法,几乎不可能收集足够多的代表性伪造人脸,且现有检测器难以泛化至所有篡改类型。因此,我们转向学习真实人脸的分布,并间接识别偏离真实人脸分布的伪造图像。在本研究中,我们提出真实人脸基础表征学习(RFFR),旨在从大规模真实人脸数据集中学习通用表征,并检测RFFR分布范围外的潜在伪影。具体而言,我们通过掩码图像建模(MIM)在真实人脸数据集上训练模型,当将模型应用于伪造样本时,输入人脸与重建人脸之间会产生差异。这种差异揭示了RFFR中未包含的低层级伪影,从而更易构建对RFFR分布范围外各类潜在伪影敏感的深度伪造检测器。大量实验表明,我们的方法带来更优的泛化性能,在跨篡改评估中显著超越现有最优方法,且通过引入额外真实人脸训练RFFR具备进一步提升的潜力。