Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. However, these tasks are highly subjective and dependent on human factors, such as one's environment, attitudes, beliefs, and lived experiences. Thus, employing LLMs (which do not have such human factors) in these tasks may result in a lack of variation in data, failing to reflect the diversity of human experiences. In this paper, we examine the role of prompting LLMs with human-like personas and asking the models to answer as if they were a specific human. This is done explicitly, with exact demographics, political beliefs, and lived experiences, or implicitly via names prevalent in specific populations. The LLM personas are then evaluated via (1) subjective annotation task (e.g., detecting toxicity) and (2) a belief generation task, where both tasks are known to vary across human factors. We examine the impact of explicit vs. implicit personas and investigate which human factors LLMs recognize and respond to. Results show that LLM personas show mixed results when reproducing known human biases, but generate generally fail to demonstrate implicit biases. We conclude that LLMs lack the intrinsic cognitive mechanisms of human thought, while capturing the statistical patterns of how people speak, which may restrict their effectiveness in complex social science applications.
翻译:大语言模型(LLMs)正日益被用于以人为中心的社会科学任务中,例如数据标注、合成数据生成以及参与对话。然而,这些任务具有高度主观性,且依赖于环境、态度、信念和生活经历等人因要素。因此,在这些任务中使用LLMs(其本身不具备此类人因特征)可能导致数据缺乏变异性,无法反映人类经验的多样性。本文研究了通过赋予LLMs拟人化角色并令其以特定人类身份作答的提示方法,具体包括显性方式(明确指定人口统计特征、政治信仰和生活经历)与隐性方式(使用特定群体中常见的姓名)。我们通过以下任务评估LLM角色:(1)主观标注任务(如毒性检测);(2)信念生成任务——这两类任务均已知会因人因差异而产生变化。我们检验了显性与隐性角色的影响,并探究LLMs能识别和响应哪些人因要素。结果表明,LLM角色在复现已知人类偏见时表现参差不齐,但普遍无法展现隐性偏见。我们得出结论:LLMs缺乏人类思维的内在认知机制,仅能捕捉人类言语的统计模式,这可能限制其在复杂社会科学应用中的有效性。