Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication. We investigate this encoding using a dual-persona rewrite paradigm, prompting ten large language models (LLMs) to rewrite naturally occurring autistic discourse from either an autistic or neurotypical persona. We uncover autistic-persona rewrites diverge significantly more in lexical form and affective register than neurotypical rewrites, despite equivalent semantic similarity. Furthermore, most models collapse cross-persona generations into near-identical outputs. To uncover the mechanisms behind this generative breakdown, we introduce a multi-agent qualitative analysis framework. Our results reveal systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary are pervasive failure modes for this task that cluster by alignment strategy rather than parameter scale. Finally, our targeted comparison with autistic human annotators demonstrates that community-insider knowledge produces systematic label reversals relative to LLM classifications. Our findings indicate that current alignment training causes persona-specific generative breakdown visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve.
翻译:安全性对齐减少了显性有害输出,但无意中编码了一种经过净化的、神经规范化的边缘化沟通表征。我们采用双人格重写范式对此编码进行研究,提示十个大型语言模型(LLM)以自闭症或神经典型人格重写自然发生的自闭症话语。我们发现,尽管语义相似性相当,但自闭症人格重写在词汇形式和情感调性上的偏离程度显著大于神经典型人格重写。此外,大多数模型将跨人格生成塌缩为近乎相同的输出。为揭示这种生成性故障背后的机制,我们引入了多智能体定性分析框架。结果显示,系统性输出擦除、刻板化幻觉及任务回避型元评论是该任务中普遍存在的失败模式,且这些模式按对齐策略而非参数规模聚类。最后,我们与自闭症人类标注者的定向比较表明,社区内部知识会产生与LLM分类相对的系统性标签反转。我们的发现表明,当前的对齐训练会导致仅在定性分析中可见的人格特异性生成故障,证实了提示工程无法解决的深层表征鸿沟。