Image safety classifiers serve as a critical component of contemporary content moderation systems on the internet. However, their resilience against user-style malicious image editing remains underexplored. Such behaviors are highly prevalent in daily scenarios but difficult to fully reproduce. To explore this vulnerability, we introduce RedEdit, a novel black-box red-teaming agent that formulates photo-editing evasion as a combinatorial search problem over edit-tool sequences. It adopts a Vision-Language-Model (VLM)-based proposer to generate semantically targeted candidate edits and a Monte Carlo Tree Search (MCTS) planner to prioritize promising edit paths while backtracking from ineffective ones. Together, the proposer and planner instantiate two key capabilities of human attackers, i.e., domain knowledge and iterative backtracking, respectively, to reproduce this practical threat. Our extensive experiments on UnsafeBench reveal profound systemic vulnerabilities: fewer than two edits on average enable 76.2% of unsafe images to evade detectors, while retaining 93.0% malicious semantics, meaning that such manipulated content remains perceptually malicious to humans while easily bypassing automated moderation. We therefore appeal to the community for more attention to this overlooked practical threat.
翻译:图像安全分类器是当代互联网内容审核系统的关键组成部分。然而,其面对用户风格恶意图像编辑的鲁棒性仍未得到充分探索。此类行为在日常场景中高度普遍,却难以被完全复现。为探究这一脆弱性,我们提出RedEdit——一种新颖的黑盒红队测试代理,将照片编辑逃逸问题建模为编辑工具序列上的组合搜索问题。它采用基于视觉语言模型(VLM)的提议器生成语义定向的候选编辑,并利用蒙特卡洛树搜索(MCTS)规划器在回溯低效路径的同时优先探索有前景的编辑路径。提议器与规划器协同工作,分别实现了人类攻击者的两项核心能力(即领域知识与迭代回溯),从而复现了这一实际威胁。我们在UnsafeBench上的大量实验揭示了深层的系统性脆弱性:平均只需不到两次编辑,即可使76.2%的不安全图像逃逸检测器,同时保留93.0%的恶意语义——这意味着此类被操控的内容对人类而言仍具有感知层面的恶意性,却能轻松绕过自动化审核。因此,我们呼吁社区对这一被忽视的实际威胁给予更多关注。