Recent research has extended beyond assessing the performance of Large Language Models (LLMs) to examining their characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics. The administration of personality tests to LLMs has emerged as a noteworthy area in this context. However, the suitability of employing psychological scales, initially devised for humans, on LLMs is a matter of ongoing debate. Our study aims to determine the reliability of applying personality assessments to LLMs, explicitly investigating whether LLMs demonstrate consistent personality traits. Analyzing responses under 2,500 settings reveals that gpt-3.5-turbo shows consistency in responses to the Big Five Inventory, indicating a high degree of reliability. Furthermore, our research explores the potential of gpt-3.5-turbo to emulate diverse personalities and represent various groups, which is a capability increasingly sought after in social sciences for substituting human participants with LLMs to reduce costs. Our findings reveal that LLMs have the potential to represent different personalities with specific prompt instructions. By shedding light on the personalization of LLMs, our study endeavors to pave the way for future explorations in this field. We have made our experimental results and the corresponding code openly accessible via https://github.com/CUHK-ARISE/LLMPersonality.
翻译:近期研究已超越评估大型语言模型(LLMs)的性能,转而从心理学角度考察其特征,承认理解其行为特征的必要性。在此背景下,对LLMs进行人格测试已成为一个值得关注的领域。然而,将最初为人类设计的心理量表应用于LLMs的适宜性仍存在持续争论。本研究旨在确定LLMs人格评估的可靠性,明确探究LLMs是否展现出一致的人格特质。对2500种设置下的响应进行分析表明,gpt-3.5-turbo在大五人格量表上的响应具有一致性,显示出高度的可靠性。此外,本研究探讨了gpt-3.5-turbo模拟不同人格和代表不同群体的潜力,这一能力在社会科学中日益受到关注,以期用LLMs替代人类被试以降低成本。我们的研究发现,LLMs能够通过特定提示指令代表不同的人格特质。通过揭示LLMs的个性化特征,本研究旨在为该领域的未来探索铺平道路。我们已将实验结果及相应代码公开于https://github.com/CUHK-ARISE/LLMPersonality。