Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimization primarily in the prompt space. In other words, these methods mainly search for better attack wording or better strategy choices, but they do not search over executable code. By moving the search into code space, we can optimize not only the final attack prompt, but also the procedure that generates it, including execution flow, reusable logic, branching, and failure-driven repair. To overcome this gap, we introduce EvoSynth, an autonomous multi-agent framework that shifts the optimization space from prompts to executable code. Instead of refining prompts directly, EvoSynth employs a multi-agent system to autonomously engineer, evolve, and execute code-based attack algorithms. Crucially, it features a code-level self-correction loop, allowing it to iteratively rewrite the code-based algorithm in response to target-model feedback and failed attempts. Through extensive experiments, we demonstrate that EvoSynth achieves an 85.5\% Attack Success Rate (ASR) against highly robust models like Claude-Sonnet-4.5 and a 95.9\% average ASR across evaluated targets, while generating attacks that are significantly more diverse than those from existing methods. We release our framework to facilitate future research on evolutionary synthesis in executable code space.


翻译:自动红队测试框架在大语言模型领域已日趋复杂,但许多方法仍主要在提示空间中进行攻击优化。换言之,这些方法主要寻找更优的攻击措辞或策略选择,却未在可执行代码空间内进行搜索。通过将搜索迁移至代码空间,我们不仅可优化最终攻击提示,还能优化生成攻击提示的流程,涵盖执行流、可复用逻辑、分支结构及故障驱动修复。为克服此局限,我们提出EvoSynth——一种将优化空间从提示转向可执行代码的自主多智能体框架。EvoSynth并非直接优化提示,而是采用多智能体系统自主设计、进化并执行基于代码的攻击算法。其关键在于内置代码级自纠错循环,能根据目标模型反馈和失败尝试迭代重写基于代码的算法。通过大量实验证明,EvoSynth在面对Claude-Sonnet-4.5等高度鲁棒模型时实现了85.5%的攻击成功率(ASR),在评估目标上的平均ASR达95.9%,同时生成的攻击多样性显著优于现有方法。我们公开该框架以促进可执行代码空间中进化合成的未来研究。

0
下载
关闭预览

相关内容

《基于强化学习的自动化红队测试》
专知会员服务
8+阅读 · 7月23日
BES:让语言模型通过双向进化搜索自我改进
专知会员服务
9+阅读 · 5月30日
《用于建模系统攻击路径的强化学习环境》
专知会员服务
23+阅读 · 3月5日
《攻击场景描述形式化模型研究》
专知会员服务
33+阅读 · 2025年8月15日
大语言模型越狱攻击:模型、根因及其攻防演化
专知会员服务
22+阅读 · 2025年4月28日
大语言模型越狱攻击: 模型、根因及其攻防演化
专知会员服务
24+阅读 · 2025年2月16日
模型攻击:鲁棒性联邦学习研究的最新进展
机器之心
35+阅读 · 2020年6月3日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
19+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
VIP会员
最新内容
《履带式无人地面战车技术发展现状》
专知会员服务
3+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
3+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
12+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
10+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
6+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
9+阅读 · 7月31日
相关VIP内容
《基于强化学习的自动化红队测试》
专知会员服务
8+阅读 · 7月23日
BES:让语言模型通过双向进化搜索自我改进
专知会员服务
9+阅读 · 5月30日
《用于建模系统攻击路径的强化学习环境》
专知会员服务
23+阅读 · 3月5日
《攻击场景描述形式化模型研究》
专知会员服务
33+阅读 · 2025年8月15日
大语言模型越狱攻击:模型、根因及其攻防演化
专知会员服务
22+阅读 · 2025年4月28日
大语言模型越狱攻击: 模型、根因及其攻防演化
专知会员服务
24+阅读 · 2025年2月16日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
19+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员