The leading strategy for analyzing unstructured data uses two steps. First, latent variables of economic interest are estimated with an upstream information retrieval model. Second, the estimates are treated as "data" in a downstream econometric model. We establish theoretical arguments for why this two-step strategy leads to biased inference in empirically plausible settings. More constructively, we propose a one-step strategy for valid inference that uses the upstream and downstream models jointly. The one-step strategy (i) substantially reduces bias in simulations; (ii) has quantitatively important effects in a leading application using CEO time-use data; and (iii) can be readily adapted by applied researchers.
翻译:分析非结构化数据的主流策略采用两步法:首先,通过上游信息检索模型估计具有经济意义的潜在变量;其次,将这些估计值作为“数据”用于下游计量经济学模型。我们建立了理论依据,证明这种两步法在实证合理的情境下会导致有偏推断。更富建设性的是,我们提出了一种联合运用上下游模型的有效推断单步策略。该单步策略:(i) 在模拟实验中显著降低了偏误;(ii) 在一项基于CEO时间使用数据的领先应用中产生了定量显著影响;(iii) 可被应用研究人员便捷采纳。