How can we use large language models (LLMs) to augment surveys? This paper investigates three distinct applications of LLMs fine-tuned by nationally representative surveys for opinion prediction -- missing data imputation, retrodiction, and zero-shot prediction. We present a new methodological framework that incorporates neural embeddings of survey questions, individual beliefs, and temporal contexts to personalize LLMs in opinion prediction. Among 3,110 binarized opinions from 68,846 Americans in the General Social Survey from 1972 to 2021, our best models based on Alpaca-7b excels in missing data imputation (AUC = 0.87 for personal opinion prediction and $\rho$ = 0.99 for public opinion prediction) and retrodiction (AUC = 0.86, $\rho$ = 0.98). These remarkable prediction capabilities allow us to fill in missing trends with high confidence and pinpoint when public attitudes changed, such as the rising support for same-sex marriage. However, the models show limited performance in a zero-shot prediction task (AUC = 0.73, $\rho$ = 0.67), highlighting challenges presented by LLMs without human responses. Further, we find that the best models' accuracy is lower for individuals with low socioeconomic status, racial minorities, and non-partisan affiliations but higher for ideologically sorted opinions in contemporary periods. We discuss practical constraints, socio-demographic representation, and ethical concerns regarding individual autonomy and privacy when using LLMs for opinion prediction. This paper showcases a new approach for leveraging LLMs to enhance nationally representative surveys by predicting missing responses and trends.
翻译:如何利用大型语言模型(LLMs)增强调查?本文研究了三种由全国代表性调查微调的LLMs在意见预测中的不同应用——缺失数据插补、回溯预测和零样本预测。我们提出了一种新的方法论框架,该框架整合了调查问题的神经嵌入、个人信念和时间背景,以在意见预测中个性化LLMs。在1972年至2021年综合社会调查中来自68,846名美国人的3,110个二值化意见中,我们基于Alpaca-7b的最佳模型在缺失数据插补(个人意见预测AUC=0.87,公众意见预测ρ=0.99)和回溯预测(AUC=0.86,ρ=0.98)中表现出色。这些显著的预测能力使我们能够高置信度地填补缺失趋势,并精确识别公众态度变化的转折点,例如对同性婚姻支持率的上升。然而,模型在零样本预测任务中表现有限(AUC=0.73,ρ=0.67),凸显了缺乏人类响应的LLMs所面临的挑战。此外,我们发现最佳模型对低社会经济地位个体、少数族裔和无党派倾向人群的准确性较低,但在当代时期对意识形态排序明确的意见预测准确性较高。我们讨论了在使用LLMs进行意见预测时的实际约束、社会人口代表性,以及关于个人自主权和隐私的伦理问题。本文展示了一种利用LLMs增强全国代表性调查的新方法,通过预测缺失响应和趋势来弥补数据不足。