Large Language Models have vastly grown in capabilities. One potential application of such AI systems is to support data collection in the social sciences, where perfect experimental control is currently unfeasible and the collection of large, representative datasets is generally expensive. In this paper, we re-replicate 14 studies from the Many Labs 2 replication project (Klein et al., 2018) with OpenAI's text-davinci-003 model, colloquially known as GPT3.5. For the 10 studies that we could analyse, we collected a total of 10,136 responses, each of which was obtained by running GPT3.5 with the corresponding study's survey inputted as text. We find that our GPT3.5-based sample replicates 30% of the original results as well as 30% of the Many Labs 2 results, although there is heterogeneity in both these numbers (as we replicate some original findings that Many Labs 2 did not and vice versa). We also find that unlike the corresponding human subjects, GPT3.5 answered some survey questions with extreme homogeneity$\unicode{x2013}$with zero variation in different runs' responses$\unicode{x2013}$raising concerns that a hypothetical AI-led future may in certain ways be subject to a diminished diversity of thought. Overall, while our results suggest that Large Language Model psychology studies are feasible, their findings should not be assumed to straightforwardly generalise to the human case. Nevertheless, AI-based data collection may eventually become a viable and economically relevant method in the empirical social sciences, making the understanding of its capabilities and applications central.
翻译:大型语言模型的能力已大幅提升。此类人工智能系统的一个潜在应用是支持社会科学中的数据收集——目前在该领域,完美的实验控制尚不可行,且收集大规模代表性数据集通常成本高昂。在本文中,我们使用OpenAI的text-davinci-003模型(俗称GPT3.5)对Many Labs 2重复实验项目(Klein等人,2018)中的14项研究进行了再重复。对于可分析的10项研究,我们共收集了10,136个回应,每个回应均通过将相应研究的问卷以文本形式输入GPT3.5获得。我们发现,基于GPT3.5的样本重复了30%的原始结果以及30%的Many Labs 2结果,尽管这两个比例存在异质性(例如,我们重复了部分Many Labs 2未能重复的原始发现,反之亦然)。我们还发现,与相应的人类受试者不同,GPT3.5对某些问卷问题的回答具有极端同质性——不同运行下的回应毫无变异——这引发了担忧:假想的AI主导未来可能在某种程度上受到思维多样性减少的影响。总体而言,尽管我们的结果表明基于大型语言模型的心理学研究是可行的,但其发现不应被假定能直接推广到人类案例。尽管如此,基于AI的数据收集最终可能成为实证社会科学中一种可行且经济相关的方法,因此理解其能力与应用至关重要。