Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This paper introduces Transient Turn Injection(TTI), a new multi-turn attack technique that systematically exploits stateless moderation by distributing adversarial intent across isolated interactions. TTI leverages automated attacker agents powered by large language models to iteratively test and evade policy enforcement in both commercial and open-source LLMs, marking a departure from conventional jailbreak approaches that typically depend on maintaining persistent conversational context. Our extensive evaluation across state-of-the-art models-including those from OpenAI, Anthropic, Google Gemini, Meta, and prominent open-source alternatives-uncovers significant variations in resilience to TTI attacks, with only select architectures exhibiting substantial inherent robustness. Our automated blackbox evaluation framework also uncovers previously unknown model specific vulnerabilities and attack surface patterns, especially within medical and high stakes domains. We further compare TTI against established adversarial prompting methods and detail practical mitigation strategies, such as session level context aggregation and deep alignment approaches. Our study underscores the urgent need for holistic, context aware defenses and continuous adversarial testing to future proof LLM deployments against evolving multi-turn threats.


翻译:大语言模型(LLM)日益融入敏感工作流程,提高了对抗鲁棒性与安全性的要求。本文提出瞬态轮次注入(TTI),一种新型多轮攻击技术,通过将对抗意图分布到孤立交互中,系统性地利用无状态审查机制。TTI 利用由大语言模型驱动的自动化攻击智能体,在商业和开源大语言模型中迭代测试并规避策略执行,这标志着与传统越狱方法(通常依赖于维持持续对话上下文)的背离。我们在包括 OpenAI、Anthropic、Google Gemini、Meta 及著名开源替代方案在内的最先进模型上的广泛评估,揭示了这些模型对 TTI 攻击鲁棒性的显著差异,仅少数架构表现出显著的内在稳健性。我们的自动化黑盒评估框架还发现了先前未知的特定模型漏洞和攻击面模式,尤其是在医疗和高风险领域。我们进一步将 TTI 与已建立的对抗提示方法进行比较,并详细介绍了实用缓解策略,例如会话级上下文聚合和深度对齐方法。我们的研究强调了为应对不断演变的多轮威胁,对 LLM 部署进行整体性、上下文感知防御和持续对抗测试的迫切需求。

0
下载
关闭预览

相关内容

多模态大语言模型的自我改进:综述
专知会员服务
29+阅读 · 2025年10月8日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
多模态大语言模型
专知会员服务
98+阅读 · 2024年6月25日
《多模态大型语言模型》最新进展,详述26种现有MM-LLMs
专知会员服务
65+阅读 · 2024年1月25日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
深度学习的下一步:Transformer和注意力机制
云头条
56+阅读 · 2019年9月14日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Instruction Tuning for Large Language Models: A Survey
Arxiv
15+阅读 · 2023年8月21日
Arxiv
25+阅读 · 2023年6月23日
VIP会员
最新内容
《美军水下战与海床战概述及本地实施》
专知会员服务
0+阅读 · 44分钟前
面向未来冲突推进陆军情报体制改革
专知会员服务
0+阅读 · 今天4:12
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
2+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
相关VIP内容
多模态大语言模型的自我改进:综述
专知会员服务
29+阅读 · 2025年10月8日
TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
多模态大语言模型
专知会员服务
98+阅读 · 2024年6月25日
《多模态大型语言模型》最新进展,详述26种现有MM-LLMs
专知会员服务
65+阅读 · 2024年1月25日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员