Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent work uses these ``digital twin'' predictions to supplement human responses in surveys and experiments. We study the special case of using AI-generated predictions to reduce variance in randomized experiments. We argue that doing so requires no new estimators and that researchers can simply include AI predictions as covariates in standard regression adjustment, analogous to adjusting for a prognostic score. A benefit of this approach is a ``do no harm'' property whereby the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative. Other methods, such as variants of prediction-powered inference, do not have this guarantee. We provide implementation guidance, including how to obtain continuous scores from discrete LLM outputs and how to use LLMs to featurize unstructured inputs as auxiliary covariates. We demonstrate these ideas in simulations and three empirical applications: a survey mega-study, an email marketing A/B test, and a large-scale technology platform experiment. Overall, efficiency gains are real if modest, with greater benefits in studies that contain substantial text and other unstructured data. We also confirm the do no harm property empirically. Given these gains and limited costs, we recommend adjusting for AI-generated predictions as a regular empirical practice.
翻译:生成式人工智能和大语言模型能够从丰富、非结构化的输入中,利用极少甚至没有任务特定训练数据,产生人类行为的现实预测。近期研究使用这些“数字孪生”预测来补充调查和实验中的人类回应。我们研究了利用AI生成预测来减少随机实验方差的特殊情形。我们主张,这样做无需新的估计量,研究者仅需将AI预测作为协变量纳入标准回归调整中,类似于对预后评分进行调整。这一方法的一个优势是具有“无害”特性:当预测不具信息量时,调整后的估计量会回退至未经调整的均值差。其他方法(如预测驱动推理的变体)则不保证此特性。我们提供了实施指南,包括如何从离散的大语言模型输出中获取连续得分,以及如何利用大语言模型将非结构化输入特征化作为辅助协变量。我们通过模拟和三个实证应用(一项大规模调查、一个邮件营销A/B测试以及一个大规模技术平台实验)展示了这些思想。总体而言,效率提升虽然温和但切实存在,在包含大量文本及其他非结构化数据的研究中收益更大。我们还从实证角度验证了“无害”特性。鉴于这些收益及有限成本,我们建议将AI生成预测的调整作为常规实证实践。