As large language models (LLMs) become more capable, there is growing excitement about the possibility of using LLMs as proxies for humans in real-world tasks where subjective labels are desired, such as in surveys and opinion polling. One widely-cited barrier to the adoption of LLMs is their sensitivity to prompt wording - but interestingly, humans also display sensitivities to instruction changes in the form of response biases. As such, we argue that if LLMs are going to be used to approximate human opinions, it is necessary to investigate the extent to which LLMs also reflect human response biases, if at all. In this work, we use survey design as a case study, where human response biases caused by permutations in wordings of "prompts" have been extensively studied. Drawing from prior work in social psychology, we design a dataset and propose a framework to evaluate whether LLMs exhibit human-like response biases in survey questionnaires. Our comprehensive evaluation of nine models shows that popular open and commercial LLMs generally fail to reflect human-like behavior. These inconsistencies tend to be more prominent in models that have been instruction fine-tuned. Furthermore, even if a model shows a significant change in the same direction as humans, we find that perturbations that are not meant to elicit significant changes in humans may also result in a similar change. These results highlight the potential pitfalls of using LLMs to substitute humans in parts of the annotation pipeline, and further underscore the importance of finer-grained characterizations of model behavior. Our code, dataset, and collected samples are available at https://github.com/lindiatjuatja/BiasMonkey
翻译:随着大语言模型能力的提升,利用其替代人类完成需要主观标注的现实任务(如问卷调查和民意测验)的可能性日益受到关注。一个被广泛提及的阻碍LLMs应用的因素是其对提示措辞的敏感性——但有趣的是,人类同样会因指令变化而产生回答偏差。基于此,我们认为若要将LLMs用于近似人类意见,有必要研究LLMs在多大程度上也反映人类的回答偏差。本研究以问卷设计为案例,其中因"提示"措辞变化导致的类人回答偏差已被广泛研究。借鉴社会心理学既有成果,我们构建了一个数据集并提出评估框架,检验LLMs在问卷中是否表现出类人回答偏差。对九个模型的全面评估表明,主流开源和商业LLMs普遍未能展现类人行为。这种不一致性在指令微调后的模型中尤为显著。此外,即使模型在相同方向上出现显著变化,那些预期不会引发人类显著变化的扰动也可能导致类似效应。这些结果揭示了在标注流程中使用LLMs替代人类的潜在缺陷,进一步凸显了对模型行为进行细粒度描述的重要性。我们的代码、数据集及采集样本已发布于https://github.com/lindiatjuatja/BiasMonkey