Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication. We investigate this encoding using a dual-persona rewrite paradigm, prompting ten large language models (LLMs) to rewrite naturally occurring autistic discourse from either an autistic or neurotypical persona. We uncover autistic-persona rewrites diverge significantly more in lexical form and affective register than neurotypical rewrites, despite equivalent semantic similarity. Furthermore, most models collapse cross-persona generations into near-identical outputs. To uncover the mechanisms behind this generative breakdown, we introduce a multi-agent qualitative analysis framework. Our results reveal systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary are pervasive failure modes for this task that cluster by alignment strategy rather than parameter scale. Finally, our targeted comparison with autistic human annotators demonstrates that community-insider knowledge produces systematic label reversals relative to LLM classifications. Our findings indicate that current alignment training causes persona-specific generative breakdown visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve.


翻译:安全性对齐减少了显性有害输出,但无意中编码了一种经过净化的、神经规范化的边缘化沟通表征。我们采用双人格重写范式对此编码进行研究,提示十个大型语言模型(LLM)以自闭症或神经典型人格重写自然发生的自闭症话语。我们发现,尽管语义相似性相当,但自闭症人格重写在词汇形式和情感调性上的偏离程度显著大于神经典型人格重写。此外,大多数模型将跨人格生成塌缩为近乎相同的输出。为揭示这种生成性故障背后的机制,我们引入了多智能体定性分析框架。结果显示,系统性输出擦除、刻板化幻觉及任务回避型元评论是该任务中普遍存在的失败模式,且这些模式按对齐策略而非参数规模聚类。最后,我们与自闭症人类标注者的定向比较表明,社区内部知识会产生与LLM分类相对的系统性标签反转。我们的发现表明,当前的对齐训练会导致仅在定性分析中可见的人格特异性生成故障,证实了提示工程无法解决的深层表征鸿沟。

0
下载
关闭预览

相关内容

大型语言模型中隐性与显性偏见的综合研究
专知会员服务
17+阅读 · 2025年11月25日
可信赖LLM智能体的研究综述:威胁与应对措施
专知会员服务
36+阅读 · 2025年3月17日
《以人为中心的大型语言模型(LLM)研究综述》
专知会员服务
41+阅读 · 2024年11月25日
【论文笔记】基于强化学习的人机对话
专知
20+阅读 · 2019年9月21日
NLP 与 NLU:从语言理解到语言处理
AI研习社
15+阅读 · 2019年5月29日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
3+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
6+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
5+阅读 · 8月5日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员