Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. We perform input-level ablation studies to assess the contribution of individual and combinations of poetic devices; construct an interpretable vector representation of attention maps; cluster these representations and train linear probes to predict safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study.


翻译:大型语言模型日益部署于安全关键应用,其抵御有害指令的能力至关重要。尽管后训练旨在使模型对多种越狱策略具有鲁棒性,但最新证据表明,风格化改写(如诗歌化转换)仍能以惊人的有效性绕过安全机制。这引发了一个核心问题:为何文学性越狱能够成功?在本研究中,我们探讨其有效性是否依赖于特定诗歌手法、是否源于对文学格式的识别失败,或是否由模型处理风格不规则提示时更深的改变所致。我们通过注意力模式的归因分析来解决该问题。具体地,我们进行输入级消融研究以评估单个及组合诗歌手法的贡献;构建注意力图的可解释向量表示;聚类这些表示并训练线性探测以预测安全结果与文学格式。结果表明,模型能高精度区分诗歌与散文格式,但在各格式内预测越狱成功的能力仍显不足。进一步聚类显示,模型按文学格式而非安全标签存在清晰分离。这些发现表明越狱成功并非源于无法识别诗歌格式;相反,诗歌提示诱发独特的处理模式,该模式在很大程度上独立于有害内容检测。总体而言,文学性越狱并非通过单一诗歌手法,而是通过累积的风格不规则性改变提示处理过程,从而规避后训练中考虑的词汇触发条件,导致大型语言模型出现不匹配。这提示鲁棒性要求安全机制能够解释风格引发的模型行为偏移。本文以 Qwen3-14B 作为代表性开源权重案例展开研究。

0
下载
关闭预览

相关内容

【ICML2026】大型视觉语言模型在注意力中迷失
专知会员服务
10+阅读 · 5月10日
大语言模型机器遗忘综述
专知会员服务
18+阅读 · 2025年11月2日
大语言模型越狱攻击:模型、根因及其攻防演化
专知会员服务
22+阅读 · 2025年4月28日
深度学习的下一步:Transformer和注意力机制
云头条
56+阅读 · 2019年9月14日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
注意力能提高模型可解释性?实验表明:并没有
黑龙江大学自然语言处理实验室
11+阅读 · 2019年4月16日
深入理解BERT Transformer ,不仅仅是注意力机制
大数据文摘
22+阅读 · 2019年3月19日
一文读懂「Attention is All You Need」| 附代码实现
PaperWeekly
37+阅读 · 2018年1月10日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关基金
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员