Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. We perform input-level ablation studies to assess the contribution of individual and combinations of poetic devices; construct an interpretable vector representation of attention maps; cluster these representations and train linear probes to predict safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study.
翻译:大型语言模型日益部署于安全关键应用,其抵御有害指令的能力至关重要。尽管后训练旨在使模型对多种越狱策略具有鲁棒性,但最新证据表明,风格化改写(如诗歌化转换)仍能以惊人的有效性绕过安全机制。这引发了一个核心问题:为何文学性越狱能够成功?在本研究中,我们探讨其有效性是否依赖于特定诗歌手法、是否源于对文学格式的识别失败,或是否由模型处理风格不规则提示时更深的改变所致。我们通过注意力模式的归因分析来解决该问题。具体地,我们进行输入级消融研究以评估单个及组合诗歌手法的贡献;构建注意力图的可解释向量表示;聚类这些表示并训练线性探测以预测安全结果与文学格式。结果表明,模型能高精度区分诗歌与散文格式,但在各格式内预测越狱成功的能力仍显不足。进一步聚类显示,模型按文学格式而非安全标签存在清晰分离。这些发现表明越狱成功并非源于无法识别诗歌格式;相反,诗歌提示诱发独特的处理模式,该模式在很大程度上独立于有害内容检测。总体而言,文学性越狱并非通过单一诗歌手法,而是通过累积的风格不规则性改变提示处理过程,从而规避后训练中考虑的词汇触发条件,导致大型语言模型出现不匹配。这提示鲁棒性要求安全机制能够解释风格引发的模型行为偏移。本文以 Qwen3-14B 作为代表性开源权重案例展开研究。