New advancements for the detection of synthetic images are critical for fighting disinformation, as the capabilities of generative AI models continuously evolve and can lead to hyper-realistic synthetic imagery at unprecedented scale and speed. In this paper, we focus on the challenge of generalizing across different concept classes, e.g., when training a detector on human faces and testing on synthetic animal images - highlighting the ineffectiveness of existing approaches that randomly sample generated images to train their models. By contrast, we propose an approach based on the premise that the robustness of the detector can be enhanced by training it on realistic synthetic images that are selected based on their quality scores according to a probabilistic quality estimation model. We demonstrate the effectiveness of the proposed approach by conducting experiments with generated images from two seminal architectures, StyleGAN2 and Latent Diffusion, and using three different concepts for each, so as to measure the cross-concept generalization ability. Our results show that our quality-based sampling method leads to higher detection performance for nearly all concepts, improving the overall effectiveness of the synthetic image detectors.
翻译:针对生成式AI模型能力的持续演进及其以前所未有的规模和速度生成超逼真合成图像的现实,新型合成图像检测技术对于打击虚假信息至关重要。本文聚焦于跨不同概念类别的泛化挑战——例如使用人脸图像训练检测器,却在动物合成图像上测试——揭示了现有方法在随机采样生成图像训练模型时存在的低效性。与此相反,我们提出一种基于以下前提的新方法:通过训练检测器处理经概率质量估计模型筛选的高质量合成图像,可提升其鲁棒性。为验证所提方法的有效性,我们采用两种经典架构(StyleGAN2和潜扩散模型)生成的图像开展实验,每种架构均涉及三种不同概念,以衡量跨概念泛化能力。实验结果表明,基于质量的采样方法在几乎所有概念场景下均能提升检测性能,显著增强了合成图像检测器的整体效果。