The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what extent specific models exhibit or promote authoritarian attitudes and characteristics. We introduce AuAu, a comprehensive benchmark that aims to assess the risk of LLMs generating responses with authoritarian tendencies. This benchmark combines three evaluation approaches: (i) psychometric questions from an extensive pool of 15 human validated instruments; (ii) contextual behavior vignettes probing intended actions in concrete situations; and (iii) responses to realistic user prompts. Unlike prior work, AuAu evaluates not only a general closeness towards authoritarianism but also the established sub-concepts Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find that all tested models exhibit substantial authoritarian response rates under the psychometric evaluation, though rates drop significantly in increasingly more realistic downstream task. We further find that an authoritarian system prompt easily manipulates 15 out of 17 models to promote increased authoritarianism. Our results underscore the need for continued, systematic auditing of LLM-based AI systems to detect and ultimately mitigate undesired authoritarian tendencies in generated output. Our code and data are available at: https://github.com/andreaseinwiller/AuAu
翻译:全球威权主义浪潮的兴起,结合其在用户日常生活中的核心地位日益增强,引发了一个问题:特定模型在何种程度上展现或宣扬威权态度与特征。我们提出AuAu,一个旨在评估大语言模型生成具有威权倾向回复风险的综合基准。该基准融合三种评估方法:(i) 来自15个经人工验证工具的大规模题库的心理测量问题;(ii) 探究具体情境下意图行为的语境行为片段;以及(iii) 对现实用户提示的响应。与既往研究不同,AuAu不仅评估对威权主义的一般亲近程度,还评估其已确立的子概念:威权攻击、威权服从与传统主义。通过评估来自中国、欧盟、俄罗斯和美国的17个模型,我们发现所有受测模型在心理测量评估下均展现出显著的威权回复率,尽管在逐渐趋近现实的下游任务中该比率显著下降。我们进一步发现,威权系统提示轻易操纵了17个模型中的15个,使其宣扬更强的威权主义。我们的结果强调了持续、系统性地审计基于大语言模型的人工智能系统的必要性,以检测并最终减轻生成输出中不良的威权倾向。我们的代码与数据发布于:https://github.com/andreaseinwiller/AuAu