The increasing deployment of large language models (LLMs) in safety-critical applications raises fundamental challenges in systematically evaluating robustness against adversarial behaviors. Existing red-teaming practices are largely manual and expert-driven, which limits scalability, reproducibility, and coverage in high-dimensional prompt spaces. We formulate automated LLM red-teaming as a structured adversarial search problem and propose a learning-driven framework for scalable vulnerability discovery. The approach combines meta-prompt-guided adversarial prompt generation with a hierarchical execution and detection pipeline, enabling standardized evaluation across six representative threat categories, including reward hacking, deceptive alignment, data exfiltration, sandbagging, inappropriate tool use, and chain-of-thought manipulation. Extensive experiments on GPT-OSS-20B identify 47 vulnerabilities, including 21 high-severity failures and 12 previously undocumented attack patterns. Compared with manual red-teaming under matched query budgets, our method achieves a 3.9$\times$ higher discovery rate with 89\% detection accuracy, demonstrating superior coverage, efficiency, and reproducibility for large-scale robustness evaluation.
翻译:大语言模型在安全关键应用中的日益部署,对系统性地评估其对抗行为的鲁棒性提出了根本性挑战。现有红队测试实践主要依赖人工和专家驱动,在高维提示空间中限制了可扩展性、可复现性和覆盖范围。本文将自动化大语言模型红队测试形式化为结构化的对抗搜索问题,并提出一种学习驱动的可扩展漏洞发现框架。该方法结合了元提示引导的对抗提示生成与分层执行及检测管道,能够对六种代表性威胁类别(包括奖励操纵、欺骗性对齐、数据窃取、沙袋效应、不当工具使用和思维链操控)进行标准化评估。在GPT-OSS-20B上的大量实验识别出47个漏洞,其中包括21个高严重性故障和12个此前未记录的对抗攻击模式。在与匹配查询预算的手动红队测试相比,我们的方法以89%的检测准确率实现了3.9倍的漏洞发现率,在大规模鲁棒性评估中展现了卓越的覆盖范围、效率和可复现性。