Image safety classifiers serve as a critical component of contemporary content moderation systems on the internet. However, their resilience against user-style malicious image editing remains underexplored. Such behaviors are highly prevalent in daily scenarios but difficult to fully reproduce. To explore this vulnerability, we introduce RedEdit, a novel black-box red-teaming agent that formulates photo-editing evasion as a combinatorial search problem over edit-tool sequences. It adopts a Vision-Language-Model (VLM)-based proposer to generate semantically targeted candidate edits and a Monte Carlo Tree Search (MCTS) planner to prioritize promising edit paths while backtracking from ineffective ones. Together, the proposer and planner instantiate two key capabilities of human attackers, i.e., domain knowledge and iterative backtracking, respectively, to reproduce this practical threat. Our extensive experiments on UnsafeBench reveal profound systemic vulnerabilities: fewer than two edits on average enable 76.2% of unsafe images to evade detectors, while retaining 93.0% malicious semantics, meaning that such manipulated content remains perceptually malicious to humans while easily bypassing automated moderation. We therefore appeal to the community for more attention to this overlooked practical threat.


翻译:图像安全分类器是当代互联网内容审核系统的关键组成部分。然而,其面对用户风格恶意图像编辑的鲁棒性仍未得到充分探索。此类行为在日常场景中高度普遍,却难以被完全复现。为探究这一脆弱性,我们提出RedEdit——一种新颖的黑盒红队测试代理,将照片编辑逃逸问题建模为编辑工具序列上的组合搜索问题。它采用基于视觉语言模型(VLM)的提议器生成语义定向的候选编辑,并利用蒙特卡洛树搜索(MCTS)规划器在回溯低效路径的同时优先探索有前景的编辑路径。提议器与规划器协同工作,分别实现了人类攻击者的两项核心能力(即领域知识与迭代回溯),从而复现了这一实际威胁。我们在UnsafeBench上的大量实验揭示了深层的系统性脆弱性:平均只需不到两次编辑,即可使76.2%的不安全图像逃逸检测器,同时保留93.0%的恶意语义——这意味着此类被操控的内容对人类而言仍具有感知层面的恶意性,却能轻松绕过自动化审核。因此,我们呼吁社区对这一被忽视的实际威胁给予更多关注。

0
下载
关闭预览

相关内容

《基于强化学习的自动化红队测试》
专知会员服务
8+阅读 · 7月23日
《大语言模型驱动的智能红队测试》
专知会员服务
18+阅读 · 2025年11月26日
《人工智能红队测试的再审视》
专知会员服务
16+阅读 · 2025年9月2日
《评估生成式人工智能的红队方法》最新37页长综述
专知会员服务
57+阅读 · 2024年5月27日
Transformer 驱动的图像分类研究进展综述
专知会员服务
56+阅读 · 2023年2月24日
面向图像分类的对抗鲁棒性评估综述
专知会员服务
59+阅读 · 2022年10月15日
专知会员服务
65+阅读 · 2020年9月10日
编辑推荐 | 红外弱小目标检测算法综述
中国图象图形学报
21+阅读 · 2020年10月12日
一行命令搞定图像质量评价
计算机视觉life
12+阅读 · 2019年12月31日
最全综述 | 图像目标检测
计算机视觉life
31+阅读 · 2019年6月24日
FaceNiff工具 - 适用于黑客的Android应用程序
黑白之道
151+阅读 · 2019年4月7日
NetworkMiner - 网络取证分析工具
黑白之道
16+阅读 · 2018年6月29日
基于图片内容的深度学习图片检索(一)
七月在线实验室
20+阅读 · 2017年10月1日
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Arxiv
0+阅读 · 6月14日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
1+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
《基于强化学习的自动化红队测试》
专知会员服务
8+阅读 · 7月23日
《大语言模型驱动的智能红队测试》
专知会员服务
18+阅读 · 2025年11月26日
《人工智能红队测试的再审视》
专知会员服务
16+阅读 · 2025年9月2日
《评估生成式人工智能的红队方法》最新37页长综述
专知会员服务
57+阅读 · 2024年5月27日
Transformer 驱动的图像分类研究进展综述
专知会员服务
56+阅读 · 2023年2月24日
面向图像分类的对抗鲁棒性评估综述
专知会员服务
59+阅读 · 2022年10月15日
专知会员服务
65+阅读 · 2020年9月10日
相关资讯
编辑推荐 | 红外弱小目标检测算法综述
中国图象图形学报
21+阅读 · 2020年10月12日
一行命令搞定图像质量评价
计算机视觉life
12+阅读 · 2019年12月31日
最全综述 | 图像目标检测
计算机视觉life
31+阅读 · 2019年6月24日
FaceNiff工具 - 适用于黑客的Android应用程序
黑白之道
151+阅读 · 2019年4月7日
NetworkMiner - 网络取证分析工具
黑白之道
16+阅读 · 2018年6月29日
基于图片内容的深度学习图片检索(一)
七月在线实验室
20+阅读 · 2017年10月1日
相关基金
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员