With large language models (LLMs) appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate various properties of these models using tests originally designed for humans. While re-using existing tests is a resource-efficient way to evaluate LLMs, careful adjustments are usually required to ensure that test results are even valid across human sub-populations. Thus, it is not clear to what extent different tests' validity generalizes to LLMs. In this work, we provide evidence that LLMs' responses to personality tests systematically deviate from typical human responses, implying that these results cannot be interpreted in the same way as human test results. Concretely, reverse-coded items (e.g. "I am introverted" vs "I am extraverted") are often both answered affirmatively by LLMs. In addition, variation across different prompts designed to "steer" LLMs to simulate particular personality types does not follow the clear separation into five independent personality factors from human samples. In light of these results, we believe it is important to pay more attention to tests' validity for LLMs before drawing strong conclusions about potentially ill-defined concepts like LLMs' "personality".
翻译:随着大语言模型在基于文本的交互中展现出越来越类人的行为,人们开始普遍尝试使用原本为人类设计的测试来评估这些模型的各类属性。尽管复用现有测试是评估大语言模型的一种资源高效的方式,但通常需要仔细调整以确保测试结果在人类亚群体之间仍然有效。因此,不同测试的有效性在多大程度上能推广至大语言模型尚不明确。本研究提供了证据表明,大语言模型对人格测试的响应系统性地偏离了典型人类响应,这意味着这些结果无法以与人类测试结果相同的方式解读。具体而言,反向编码项目(例如,“我是内向的”与“我是外向的”)常被大语言模型同时肯定回答。此外,针对旨在“引导”大语言模型模拟特定人格类型的不同提示所产生的变异,并未遵循人类样本中五个独立人格因素的清晰分离。基于这些结果,我们认为,在就大语言模型的“人格”等可能定义不清的概念得出强有力结论之前,更应关注测试对大语言模型的有效性。