In the era of rapid digitalization and artificial intelligence advancements, the development of DeepFake technology has posed significant security and privacy concerns. This paper presents an effective measure to assess the visual realism of DeepFake videos. We utilize an ensemble of two Convolutional Neural Network (CNN) models: Eva and ConvNext. These models have been trained on the DeepFake Game Competition (DFGC) 2022 dataset and aim to predict Mean Opinion Scores (MOS) from DeepFake videos based on features extracted from sequences of frames. Our method secured the third place in the recent DFGC on Visual Realism Assessment held in conjunction with the 2023 International Joint Conference on Biometrics (IJCB 2023). We provide an over\-view of the models, data preprocessing, and training procedures. We also report the performance of our models against the competition's baseline model and discuss the implications of our findings.
翻译:在快速数字化和人工智能发展的时代,深度伪造技术的进步已引发显著的安全与隐私担忧。本文提出了一种评估深度伪造视频视觉真实性的有效方法。我们采用了两个卷积神经网络(CNN)模型的集成:Eva和ConvNext。这些模型基于DeepFake Game Competition(DFGC)2022数据集进行训练,旨在依据从帧序列中提取的特征预测深度伪造视频的均值意见得分(MOS)。我们的方法在近期与2023年国际生物识别联合会议(IJCB 2023)联合举办的DFGC视觉真实性评估比赛中荣获第三名。本文概述了模型、数据预处理及训练流程,同时报告了我们的模型与竞赛基线模型相比的性能,并讨论了研究结果的意义。