This paper present a comprehensive comparative analysis of supervised and self-supervised models for deepfake detection. We evaluate eight supervised deep learning architectures and two transformer-based models pre-trained using self-supervised strategies (DINO, CLIP) on four benchmarks (FakeAVCeleb, CelebDF-V2, DFDC, and FaceForensics++). Our analysis includes intra-dataset and inter-dataset evaluations, examining the best performing models, generalisation capabilities, and impact of augmentations. We also investigate the trade-off between model size and performance. Our main goal is to provide insights into the effectiveness of different deep learning architectures (transformers, CNNs), training strategies (supervised, self-supervised), and deepfake detection benchmarks. These insights can help guide the development of more accurate and reliable deepfake detection systems, which are crucial in mitigating the harmful impact of deepfakes on individuals and society.
翻译:本文对监督与自监督模型在深度伪造检测任务中的性能进行了全面的比较性分析。我们在四个基准数据集(FakeAVCeleb、CelebDF-V2、DFDC和FaceForensics++)上评估了八种监督深度学习架构与两种基于自监督策略预训练的Transformer模型(DINO、CLIP)。分析涵盖数据集内与跨数据集的评估,考察了最优模型性能、泛化能力及数据增强的影响。同时,我们探究了模型大小与性能之间的权衡关系。本研究旨在揭示不同深度学习架构(Transformer、CNN)、训练策略(监督、自监督)及深度伪造检测基准的有效性,为开发更精确、更可靠的深度伪造检测系统提供指导——这类系统对于减轻深度伪造对个人及社会造成的危害至关重要。