Large language models (LLMs) may not equitably represent diverse global perspectives on societal issues. In this paper, we develop a quantitative framework to evaluate whose opinions model-generated responses are more similar to. We first build a dataset, GlobalOpinionQA, comprised of questions and answers from cross-national surveys designed to capture diverse opinions on global issues across different countries. Next, we define a metric that quantifies the similarity between LLM-generated survey responses and human responses, conditioned on country. With our framework, we run three experiments on an LLM trained to be helpful, honest, and harmless with Constitutional AI. By default, LLM responses tend to be more similar to the opinions of certain populations, such as those from the USA, and some European and South American countries, highlighting the potential for biases. When we prompt the model to consider a particular country's perspective, responses shift to be more similar to the opinions of the prompted populations, but can reflect harmful cultural stereotypes. When we translate GlobalOpinionQA questions to a target language, the model's responses do not necessarily become the most similar to the opinions of speakers of those languages. We release our dataset for others to use and build on. Our data is at https://huggingface.co/datasets/Anthropic/llm_global_opinions. We also provide an interactive visualization at https://llmglobalvalues.anthropic.com.
翻译:大型语言模型(LLMs)可能无法公平地呈现全球社会议题中多元化的观点。本文构建了一个定量框架,用于评估模型生成回复与何种人群的观点更为相似。我们首先建立了GlobalOpinionQA数据集,该数据源自跨国调查中的问答对,旨在捕捉不同国家关于全球议题的多元观点。随后我们定义了基于国家条件的量化指标,用于比较LLM生成调查回复与人类回复的相似度。基于该框架,我们对采用《宪法式AI》训练(具备"有益、诚实、无害"特性)的LLM进行了三项实验。实验表明:在默认状态下,LLM回复更倾向于与美国、部分欧洲及南美国家等特定人群的观点一致,揭示了潜在偏差;当提示模型考虑特定国家视角时,回复会趋向与提示国人群观点更相似,但可能反映有害的文化刻板印象;当将GlobalOpinionQA问题翻译为目标语言时,模型回复并不必然与相应语言使用者的观点最为契合。我们已将数据集公开发布以供学界使用和拓展:https://huggingface.co/datasets/Anthropic/llm_global_opinions,同步提供交互式可视化工具:https://llmglobalvalues.anthropic.com。