Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This paper introduces Transient Turn Injection(TTI), a new multi-turn attack technique that systematically exploits stateless moderation by distributing adversarial intent across isolated interactions. TTI leverages automated attacker agents powered by large language models to iteratively test and evade policy enforcement in both commercial and open-source LLMs, marking a departure from conventional jailbreak approaches that typically depend on maintaining persistent conversational context. Our extensive evaluation across state-of-the-art models-including those from OpenAI, Anthropic, Google Gemini, Meta, and prominent open-source alternatives-uncovers significant variations in resilience to TTI attacks, with only select architectures exhibiting substantial inherent robustness. Our automated blackbox evaluation framework also uncovers previously unknown model specific vulnerabilities and attack surface patterns, especially within medical and high stakes domains. We further compare TTI against established adversarial prompting methods and detail practical mitigation strategies, such as session level context aggregation and deep alignment approaches. Our study underscores the urgent need for holistic, context aware defenses and continuous adversarial testing to future proof LLM deployments against evolving multi-turn threats.


翻译:大语言模型(LLM)日益融入敏感工作流程,提高了对抗鲁棒性与安全性的要求。本文提出瞬态轮次注入(TTI),一种新型多轮攻击技术,通过将对抗意图分布到孤立交互中,系统性地利用无状态审查机制。TTI 利用由大语言模型驱动的自动化攻击智能体,在商业和开源大语言模型中迭代测试并规避策略执行,这标志着与传统越狱方法(通常依赖于维持持续对话上下文)的背离。我们在包括 OpenAI、Anthropic、Google Gemini、Meta 及著名开源替代方案在内的最先进模型上的广泛评估,揭示了这些模型对 TTI 攻击鲁棒性的显著差异,仅少数架构表现出显著的内在稳健性。我们的自动化黑盒评估框架还发现了先前未知的特定模型漏洞和攻击面模式,尤其是在医疗和高风险领域。我们进一步将 TTI 与已建立的对抗提示方法进行比较,并详细介绍了实用缓解策略,例如会话级上下文聚合和深度对齐方法。我们的研究强调了为应对不断演变的多轮威胁,对 LLM 部署进行整体性、上下文感知防御和持续对抗测试的迫切需求。

0
下载
关闭预览

相关内容

多模态大语言模型的自我改进:综述
专知会员服务
29+阅读 · 2025年10月8日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
多模态大语言模型
专知会员服务
98+阅读 · 2024年6月25日
《多模态大型语言模型》最新进展,详述26种现有MM-LLMs
专知会员服务
65+阅读 · 2024年1月25日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
深度学习的下一步:Transformer和注意力机制
云头条
56+阅读 · 2019年9月14日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Instruction Tuning for Large Language Models: A Survey
Arxiv
15+阅读 · 2023年8月21日
Arxiv
25+阅读 · 2023年6月23日
VIP会员
最新内容
机器的崛起:美海军陆战队组建机器人营思考
专知会员服务
5+阅读 · 9月8日
无面之战:人工智能如何重绘权力版图
专知会员服务
3+阅读 · 9月8日
《最强大的军事网状网络》
专知会员服务
9+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
7+阅读 · 9月7日
相关VIP内容
多模态大语言模型的自我改进:综述
专知会员服务
29+阅读 · 2025年10月8日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
多模态大语言模型
专知会员服务
98+阅读 · 2024年6月25日
《多模态大型语言模型》最新进展,详述26种现有MM-LLMs
专知会员服务
65+阅读 · 2024年1月25日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员