Recent advancements in language-image models have led to the development of highly realistic images that can be generated from textual descriptions. However, the increased visual quality of these generated images poses a potential threat to the field of media forensics. This paper aims to investigate the level of challenge that language-image generation models pose to media forensics. To achieve this, we propose a new approach that leverages the DALL-E2 language-image model to automatically generate and splice masked regions guided by a text prompt. To ensure the creation of realistic manipulations, we have designed an annotation platform with human checking to verify reasonable text prompts. This approach has resulted in the creation of a new image dataset called AutoSplice, containing 5,894 manipulated and authentic images. Specifically, we have generated a total of 3,621 images by locally or globally manipulating real-world image-caption pairs, which we believe will provide a valuable resource for developing generalized detection methods in this area. The dataset is evaluated under two media forensic tasks: forgery detection and localization. Our extensive experiments show that most media forensic models struggle to detect the AutoSplice dataset as an unseen manipulation. However, when fine-tuned models are used, they exhibit improved performance in both tasks.
翻译:近期语言-图像模型的进展使得能够根据文本描述生成高度逼真的图像。然而,这些生成图像视觉质量的提升对媒体取证领域构成了潜在威胁。本文旨在探究语言-图像生成模型对媒体取证带来的挑战程度。为此,我们提出了一种新方法,利用DALL-E2语言-图像模型根据文本提示自动生成并拼接掩码区域。为确保生成逼真的篡改结果,我们设计了一个带有人工审核的注释平台来验证合理的文本提示。该方法最终构建了一个名为AutoSplice的新图像数据集,包含5,894张篡改与真实图像。具体而言,我们通过局部或全局篡改真实世界的图像-文本对共生成了3,621张图像,这为开发该领域的通用检测方法提供了宝贵资源。该数据集在两项媒体取证任务(伪造检测与定位)中进行了评估。大量实验表明,多数媒体取证模型难以将AutoSplice数据集检测为未见过的篡改操作;然而,使用微调模型时,其在两项任务中均表现出更优性能。