Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This paper introduces Transient Turn Injection(TTI), a new multi-turn attack technique that systematically exploits stateless moderation by distributing adversarial intent across isolated interactions. TTI leverages automated attacker agents powered by large language models to iteratively test and evade policy enforcement in both commercial and open-source LLMs, marking a departure from conventional jailbreak approaches that typically depend on maintaining persistent conversational context. Our extensive evaluation across state-of-the-art models-including those from OpenAI, Anthropic, Google Gemini, Meta, and prominent open-source alternatives-uncovers significant variations in resilience to TTI attacks, with only select architectures exhibiting substantial inherent robustness. Our automated blackbox evaluation framework also uncovers previously unknown model specific vulnerabilities and attack surface patterns, especially within medical and high stakes domains. We further compare TTI against established adversarial prompting methods and detail practical mitigation strategies, such as session level context aggregation and deep alignment approaches. Our study underscores the urgent need for holistic, context aware defenses and continuous adversarial testing to future proof LLM deployments against evolving multi-turn threats.
翻译:大语言模型(LLM)日益融入敏感工作流程,提高了对抗鲁棒性与安全性的要求。本文提出瞬态轮次注入(TTI),一种新型多轮攻击技术,通过将对抗意图分布到孤立交互中,系统性地利用无状态审查机制。TTI 利用由大语言模型驱动的自动化攻击智能体,在商业和开源大语言模型中迭代测试并规避策略执行,这标志着与传统越狱方法(通常依赖于维持持续对话上下文)的背离。我们在包括 OpenAI、Anthropic、Google Gemini、Meta 及著名开源替代方案在内的最先进模型上的广泛评估,揭示了这些模型对 TTI 攻击鲁棒性的显著差异,仅少数架构表现出显著的内在稳健性。我们的自动化黑盒评估框架还发现了先前未知的特定模型漏洞和攻击面模式,尤其是在医疗和高风险领域。我们进一步将 TTI 与已建立的对抗提示方法进行比较,并详细介绍了实用缓解策略,例如会话级上下文聚合和深度对齐方法。我们的研究强调了为应对不断演变的多轮威胁,对 LLM 部署进行整体性、上下文感知防御和持续对抗测试的迫切需求。