A growing body of research examines personality traits in Large Language Models (LLMs), particularly in human-agent collaboration. Prior work has frequently applied the Big Five inventory to assess LLM behavior analogous to human personality, without questioning the underlying assumptions. This paper critically evaluates whether LLM responses to personality tests satisfy six defining characteristics of personality. We find that none are fully met, indicating that such assessments do not measure a construct equivalent to human personality. We propose a research agenda for shifting from anthropomorphic trait attribution toward functional evaluations, clarifying what personality tests actually capture in LLMs and developing LLM-specific frameworks for characterizing stable, intrinsic behavior.
翻译:越来越多的研究关注大型语言模型(LLMs)中的人格特质,尤其是在人机协作领域。先前的研究常采用大五人格问卷来评估与人类人格相似的LLM行为,却未质疑其基本假设。本文批判性地审视了LLM对人格测试的反应是否满足人格的六大核心特征。我们发现没有一项特征被完全满足,这表明此类评估并未测量出等同于人类人格的构念。我们提出一个研究议程,主张从拟人化特质归因转向功能性评估,澄清人格测试在LLM中实际反映的内容,并构建针对LLM的框架,以表征其稳定且内在的行为特征。