Large language models enable inexpensive AI-generated annotations, but using them reliably for causal inference remains challenging. Naively pooling AI and human data induces bias, while existing methods such as Prediction-Powered Inference (PPI; Angelopoulos et al., 2023a) treat AI outputs as proxies of true labels -- an assumption often violated for generative model outputs in practice. We propose Generative Augmented Inference (GAI), a framework that treats AI outputs as general, potentially high-dimensional informative features for learning human labels rather than as surrogates. GAI flexibly models this relationship using nonparametric methods, enabling consistent estimation and valid inference from combined human and AI data. We establish asymptotic normality and show that, under random labeling, GAI strictly improves asymptotic efficiency over human-data-only estimation whenever AI outputs are informative for true labels. Empirical studies on real-world datasets demonstrate that GAI significantly reduces estimation error and improves confidence interval quality across diverse generative data sources relative to human-only and PPI-based estimation.
翻译:大型语言模型使得低成本的人工智能生成标注变得可行,但如何可靠地将其用于因果推断仍具挑战性。简单地混合人工智能与人类数据会引入偏差,而现有方法如预测驱动推断(PPI;Angelopoulos等人,2023a)将人工智能输出视为真实标签的代理变量——这一假设在实际生成模型输出中往往不成立。我们提出生成增强推断(GAI)框架,将人工智能输出视为用于学习人类标签的通用、潜在高维信息特征,而非替代变量。GAI通过非参数方法灵活建模这种关系,实现人类与人工智能数据联合下的一致性估计与有效推断。我们建立了渐近正态性,并证明在随机标注条件下,只要人工智能输出对真实标签具有信息性,GAI严格优于仅依赖人类数据的估计在渐近效率上的表现。基于真实数据集的实证研究表明,相较于纯人类数据和基于PPI的估计,GAI能显著降低各类生成数据源的估计误差并提升置信区间质量。