Deepfakes have become a growing concern in recent years, prompting researchers to develop benchmark datasets and detection algorithms to tackle the issue. However, existing datasets suffer from significant drawbacks that hamper their effectiveness. Notably, these datasets fail to encompass the latest deepfake videos produced by state-of-the-art methods that are being shared across various platforms. This limitation impedes the ability to keep pace with the rapid evolution of generative AI techniques employed in real-world deepfake production. Our contributions in this IRB-approved study are to bridge this knowledge gap from current real-world deepfakes by providing in-depth analysis. We first present the largest and most diverse and recent deepfake dataset (RWDF-23) collected from the wild to date, consisting of 2,000 deepfake videos collected from 4 platforms targeting 4 different languages span created from 21 countries: Reddit, YouTube, TikTok, and Bilibili. By expanding the dataset's scope beyond the previous research, we capture a broader range of real-world deepfake content, reflecting the ever-evolving landscape of online platforms. Also, we conduct a comprehensive analysis encompassing various aspects of deepfakes, including creators, manipulation strategies, purposes, and real-world content production methods. This allows us to gain valuable insights into the nuances and characteristics of deepfakes in different contexts. Lastly, in addition to the video content, we also collect viewer comments and interactions, enabling us to explore the engagements of internet users with deepfake content. By considering this rich contextual information, we aim to provide a holistic understanding of the {evolving} deepfake phenomenon and its impact on online platforms.
翻译:近年来,深度伪造技术日益引发关注,促使研究人员开发基准数据集和检测算法以应对该问题。然而,现有数据集存在显著缺陷,制约了其有效性。值得注意的是,这些数据集未能涵盖当前最先进方法生成的最新深度伪造视频——这些视频正在各类平台传播。这一局限阻碍了研究人员跟上现实世界深度伪造生产中快速演进的生成式AI技术。本研究经机构审查委员会批准,旨在通过对当前真实深度伪造现象进行深入分析,填补这一知识鸿沟。我们首先提出迄今为止规模最大、最多样化且最新的野外采集深度伪造数据集(RWDF-23),包含从Reddit、YouTube、TikTok和Bilibili四个平台收集的2000个深度伪造视频,覆盖四种语言及21个原创国家。通过将数据集范围扩展至前人研究之外,我们捕捉到更广泛的现实世界深度伪造内容,反映了在线平台不断演变的态势。同时,我们开展全面分析,涵盖深度伪造的创作者、操作策略、目的及现实内容生产方式等多维度特征,从而深入理解不同情境下深度伪造的细微差异与特性。最后,除视频内容外,我们还收集观众评论与互动数据,探究互联网用户对深度伪造内容的参与模式。通过整合这些丰富的语境信息,我们旨在提供对不断演变的深度伪造现象及其对在线平台影响的整体理解。