The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what extent specific models exhibit or promote authoritarian attitudes and characteristics. We introduce AuAu, a comprehensive benchmark that aims to assess the risk of LLMs generating responses with authoritarian tendencies. This benchmark combines three evaluation approaches: (i) psychometric questions from an extensive pool of 15 human validated instruments; (ii) contextual behavior vignettes probing intended actions in concrete situations; and (iii) responses to realistic user prompts. Unlike prior work, AuAu evaluates not only a general closeness towards authoritarianism but also the established sub-concepts Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find that all tested models exhibit substantial authoritarian response rates under the psychometric evaluation, though rates drop significantly in increasingly more realistic downstream task. We further find that an authoritarian system prompt easily manipulates 15 out of 17 models to promote increased authoritarianism. Our results underscore the need for continued, systematic auditing of LLM-based AI systems to detect and ultimately mitigate undesired authoritarian tendencies in generated output. Our code and data are available at: https://github.com/andreaseinwiller/AuAu


翻译:全球威权主义浪潮的兴起,结合其在用户日常生活中的核心地位日益增强,引发了一个问题:特定模型在何种程度上展现或宣扬威权态度与特征。我们提出AuAu,一个旨在评估大语言模型生成具有威权倾向回复风险的综合基准。该基准融合三种评估方法:(i) 来自15个经人工验证工具的大规模题库的心理测量问题;(ii) 探究具体情境下意图行为的语境行为片段;以及(iii) 对现实用户提示的响应。与既往研究不同,AuAu不仅评估对威权主义的一般亲近程度,还评估其已确立的子概念:威权攻击、威权服从与传统主义。通过评估来自中国、欧盟、俄罗斯和美国的17个模型,我们发现所有受测模型在心理测量评估下均展现出显著的威权回复率,尽管在逐渐趋近现实的下游任务中该比率显著下降。我们进一步发现,威权系统提示轻易操纵了17个模型中的15个,使其宣扬更强的威权主义。我们的结果强调了持续、系统性地审计基于大语言模型的人工智能系统的必要性,以检测并最终减轻生成输出中不良的威权倾向。我们的代码与数据发布于:https://github.com/andreaseinwiller/AuAu

0
下载
关闭预览

相关内容

大语言模型基准综述
专知会员服务
27+阅读 · 2025年8月22日
大规模视觉-语言模型的基准、评估、应用与挑战
专知会员服务
18+阅读 · 2025年2月10日
VILA-U:一个融合视觉理解与生成的统一基础模型
专知会员服务
21+阅读 · 2024年9月9日
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
什么是语义角色标注?
人工智能头条
18+阅读 · 2019年4月28日
IMU 标定 | 工业界和学术界有什么不同?
计算机视觉life
14+阅读 · 2019年1月8日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
3+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
5+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
8+阅读 · 7月19日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员