Deepfake or synthetic images produced using deep generative models pose serious risks to online platforms. This has triggered several research efforts to accurately detect deepfake images, achieving excellent performance on publicly available deepfake datasets. In this work, we study 8 state-of-the-art detectors and argue that they are far from being ready for deployment due to two recent developments. First, the emergence of lightweight methods to customize large generative models, can enable an attacker to create many customized generators (to create deepfakes), thereby substantially increasing the threat surface. We show that existing defenses fail to generalize well to such \emph{user-customized generative models} that are publicly available today. We discuss new machine learning approaches based on content-agnostic features, and ensemble modeling to improve generalization performance against user-customized models. Second, the emergence of \textit{vision foundation models} -- machine learning models trained on broad data that can be easily adapted to several downstream tasks -- can be misused by attackers to craft adversarial deepfakes that can evade existing defenses. We propose a simple adversarial attack that leverages existing foundation models to craft adversarial samples \textit{without adding any adversarial noise}, through careful semantic manipulation of the image content. We highlight the vulnerabilities of several defenses against our attack, and explore directions leveraging advanced foundation models and adversarial training to defend against this new threat.
翻译:利用深度生成模型生成的深度伪造或合成图像,对在线平台构成严重威胁。这推动多项研究致力于准确检测深度伪造图像,并在公开的深度伪造数据集上取得了优异性能。本研究分析了8种最新检测技术,并认为由于两项新进展,这些技术远未达到实际部署条件。其一,轻量级定制大型生成模型方法的出现,使攻击者能够创建大量定制化生成器(用于生成深度伪造图像),从而显著扩大威胁面。研究表明,现有防御手段无法有效泛化至当前公开的此类**用户定制生成模型**。本文讨论基于内容无关特征与集成建模的新型机器学习方法,以提升针对用户定制模型的泛化性能。其二,**视觉基础模型**(在广泛数据上训练、可便捷适配多项下游任务的机器学习模型)可能被攻击者滥用,生成能规避现有防御的对抗性深度伪造。本文提出一种简单的对抗攻击方法,通过精心语义操纵图像内容,利用现有基础模型**不添加任何对抗噪声**即可生成对抗样本。研究揭示了多种防御机制对此类攻击的脆弱性,并探索了利用先进基础模型与对抗训练应对这一新型威胁的潜在方向。