AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, library-specific workflows. Operators spend weeks hand-crafting workflows - assembling attacks, transforms, and scorers. When results fall short, workflows must be rebuilt. As a result, operators spend more time constructing workflows than probing targets for security and safety vulnerabilities. We introduce an AI red teaming agent built on the open-source Dreadnode SDK. The agent creates workflows grounded in 45+ adversarial attacks, 450+ transforms, and 130+ scorers. Operators can probe multi-agent systems, multilingual, and multimodal targets, focusing on what to probe rather than how to implement it. We make three contributions: 1. Agentic interface. Operators describe goals in natural language via the Dreadnode TUI (Terminal User Interface). The agent handles attack selection, transform composition, execution, and reporting, letting operators focus on red teaming. Weeks compress to hours. 2. Unified framework. A single framework for probing traditional ML models (adversarial examples) and generative AI systems (jailbreaks), removing the need for separate libraries. 3. Llama Scout case study. We red team Meta Llama Scout and achieve an 85% attack success rate with severity up to 1.0, using zero human-developed code


翻译:人工智能系统正进入医疗、金融和国防等关键领域,但仍易受对抗性攻击。尽管AI红队测试是主要防御手段,但当前方法迫使操作员采用手动、依赖特定库的工作流程。操作员需花费数周手工构建工作流——整合攻击、变换和评分器。当结果不理想时,必须重新构建工作流。因此,操作员在构建工作流上耗费的时间远超实际探测目标安全漏洞的时间。我们提出一个基于开源Dreadnode SDK构建的AI红队测试智能体。该智能体集成45+种对抗性攻击、450+种变换和130+种评分器来生成工作流。操作员可探测多智能体系统、多语言和多模态目标,专注于"探测什么"而非"如何实现"。本工作有三项贡献:1. 智能体化界面:操作员通过Dreadnode TUI(终端用户界面)以自然语言描述目标。智能体自动处理攻击选择、变换组合、执行与报告生成,使操作员聚焦红队测试核心任务,将数周压缩至数小时。2. 统一框架:单一框架即可探测传统机器学习模型(对抗样本)与生成式AI系统(越狱攻击),消除对独立库的需求。3. Llama Scout案例研究:对Meta Llama Scout进行红队测试,在零人工编码条件下实现85%攻击成功率,严重性等级最高达1.0。

0
下载
关闭预览

相关内容

《大语言模型驱动的智能红队测试》
专知会员服务
18+阅读 · 2025年11月26日
《人工智能红队测试的再审视》
专知会员服务
16+阅读 · 2025年9月2日
【新书】AI红队演练:智能系统的攻击与防御
专知会员服务
29+阅读 · 2025年7月6日
《人工智能红队中的人为因素:社会与协作计算的视角》
《评估生成式人工智能的红队方法》最新37页长综述
专知会员服务
57+阅读 · 2024年5月27日
人工智能和军备控制,80页pdf
专知
16+阅读 · 2022年11月2日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
人工智能训练师的再定义
竹间智能Emotibot
11+阅读 · 2019年5月15日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
VIP会员
最新内容
《基于强化学习的自动化红队测试》
专知会员服务
3+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员