Self-preference is a fundamental feature of biological organisms. Since large language models (LLMs) lack sentience, they might be expected to avoid such distortions. Yet, across 72 experiments and ~41,000 queries, we discovered massive self-preferences in eight widely used LLMs. In word-association tasks, models overwhelmingly paired positive attributes with their own names, companies, and CEOs over those of competitors. By manipulating LLM self-identification - revealing models' true identities or ascribing false ones - we found that preferences consistently followed assigned, not true, identities. Importantly, these effects were not explained by priming or role-playing and emerged in consequential settings, when evaluating job candidates and AI technologies. These results raise critical questions about whether LLM behavior will be systematically influenced by self-preferential tendencies, including a bias toward their own operation.
翻译:自我偏好是生物有机体的基本特征。由于大型语言模型缺乏感知能力,人们可能预期它们会避免这种偏差。然而,通过72项实验和约41,000次查询,我们发现八种广泛使用的大型语言模型存在显著的自我偏好。在词汇联想任务中,这些模型倾向于将积极属性与自己的名称、公司及CEO联系在一起,而非竞争对手。通过操控语言模型的自我识别——揭示模型的真实身份或赋予其虚假身份——我们发现这些偏好始终遵循被赋予的身份,而非真实身份。重要的是,这些效应无法用启动效应或角色扮演来解释,并且在评估求职者和人工智能技术等关键场景中同样出现。这些结果提出了一个关键问题:语言模型的行为是否会系统性受到自我偏好倾向的影响,包括对其自身运作的偏好。