Detecting video deepfakes has become increasingly urgent in recent years. Given the audio-visual information in videos, existing methods typically expose deepfakes by modeling cross-modal correspondence using specifically designed architectures with publicly available datasets. While they have shown promising results, their effectiveness often degrades in real-world scenarios, as the limited diversity of training datasets naturally restricts generalizability to unseen cases. To address this, we propose a simple yet effective method, called AVPF, which can notably enhance model generalizability by training with self-generated Audio-Visual Pseudo-Fakes.The key idea of AVPF is to create pseudo-fake training samples that contain diverse audio-visual correspondence patterns commonly observed in real-world deepfakes. We highlight that AVPF is generated solely from authentic samples, and training relies only on authentic data and AVPF, without requiring any real deepfakes.Extensive experiments on multiple standard datasets demonstrate the strong generalizability of the proposed method, achieving an average performance improvement of up to 7.4%.
翻译:近年来,视频深度伪造检测日益迫切。现有方法通常利用公开数据集设计特定架构来建模音视频跨模态对应关系,从而识别深度伪造内容。尽管这些方法取得了显著进展,但其在真实场景中的有效性往往会因训练数据集的多样性不足而受限于对未见过案例的泛化能力。为此,我们提出一种简洁高效的方法——AVPF,通过利用自生成的音视频伪样本进行训练,显著增强模型泛化能力。AVPF的核心思想是创建包含真实世界深度伪造中常见多样化音视频对应模式的伪训练样本。我们强调,AVPF完全由真实样本生成,且训练过程仅依赖真实数据和AVPF,无需任何真实深度伪造样本。在多个标准数据集上的大量实验表明,该方法具有强大的泛化能力,平均性能提升高达7.4%。