The increasing deployment of large language models (LLMs) in safety-critical applications raises fundamental challenges in systematically evaluating robustness against adversarial behaviors. Existing red-teaming practices are largely manual and expert-driven, which limits scalability, reproducibility, and coverage in high-dimensional prompt spaces. We formulate automated LLM red-teaming as a structured adversarial search problem and propose a learning-driven framework for scalable vulnerability discovery. The approach combines meta-prompt-guided adversarial prompt generation with a hierarchical execution and detection pipeline, enabling standardized evaluation across six representative threat categories, including reward hacking, deceptive alignment, data exfiltration, sandbagging, inappropriate tool use, and chain-of-thought manipulation. Extensive experiments on GPT-OSS-20B identify 47 vulnerabilities, including 21 high-severity failures and 12 previously undocumented attack patterns. Compared with manual red-teaming under matched query budgets, our method achieves a 3.9$\times$ higher discovery rate with 89\% detection accuracy, demonstrating superior coverage, efficiency, and reproducibility for large-scale robustness evaluation.


翻译:大语言模型在安全关键应用中的日益部署,对系统性地评估其对抗行为的鲁棒性提出了根本性挑战。现有红队测试实践主要依赖人工和专家驱动,在高维提示空间中限制了可扩展性、可复现性和覆盖范围。本文将自动化大语言模型红队测试形式化为结构化的对抗搜索问题,并提出一种学习驱动的可扩展漏洞发现框架。该方法结合了元提示引导的对抗提示生成与分层执行及检测管道,能够对六种代表性威胁类别(包括奖励操纵、欺骗性对齐、数据窃取、沙袋效应、不当工具使用和思维链操控)进行标准化评估。在GPT-OSS-20B上的大量实验识别出47个漏洞,其中包括21个高严重性故障和12个此前未记录的对抗攻击模式。在与匹配查询预算的手动红队测试相比,我们的方法以89%的检测准确率实现了3.9倍的漏洞发现率,在大规模鲁棒性评估中展现了卓越的覆盖范围、效率和可复现性。

0
下载
关闭预览

相关内容

《军事大语言模型的拒绝率测量与消除》
专知会员服务
14+阅读 · 3月13日
《大语言模型驱动的智能红队测试》
专知会员服务
18+阅读 · 2025年11月26日
《人工智能红队测试的再审视》
专知会员服务
16+阅读 · 2025年9月2日
大语言模型评估技术研究进展
专知会员服务
49+阅读 · 2024年7月9日
大型语言模型在预测和异常检测中的应用综述
专知会员服务
70+阅读 · 2024年2月19日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
SemanticAdv:基于语义属性的对抗样本生成方法
机器之心
14+阅读 · 2019年7月12日
【好文解析】ICASSP最佳学生论文:深度对抗声学模型训练框架
中国科学院自动化研究所
13+阅读 · 2018年4月28日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
伊朗不对称防空战略的演进
专知会员服务
1+阅读 · 今天11:10
对抗环境下超视距目标打击的情报支援
专知会员服务
10+阅读 · 7月22日
《无人机对海面作战影响评估》
专知会员服务
15+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
7+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
8+阅读 · 7月20日
相关VIP内容
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员