In many experimental or quasi-experimental studies, outcomes of interest are only observed for subjects who select (or are selected) to engage in the activity generating the outcome. Outcome data is thus endogenously missing for units who do not engage, in which case random or conditionally random treatment assignment prior to such choices is insufficient to point identify treatment effects. Non-parametric partial identification bounds are a way to address endogenous missingness without having to make disputable parametric assumptions. Basic bounding approaches often yield bounds that are very wide and therefore minimally informative. We present methods for narrowing non-parametric bounds on treatment effects by adjusting for potentially large numbers of covariates, working with generalized random forests. Our approach allows for agnosticism about the data-generating process and honest inference. We use a simulation study and two replication exercises to demonstrate the benefits of our approach.
翻译:在许多实验或准实验研究中,感兴趣的结果仅对选择(或被选择)参与活动并产生结果的受试者可观测。对于未参与的个体,结果数据因此存在内生缺失,此时在决策前的随机或条件随机处理分配不足以点识别处理效应。非参数部分识别边界是在无需做出有争议的参数假设下解决内生缺失问题的一种方法。基本边界方法通常得到的边界过宽,因而信息量极小。我们提出通过利用广义随机森林调整大量协变量来缩窄处理效应的非参数边界。该方法允许对数据生成过程保持不可知论态度,并实现诚实推断。我们通过模拟研究和两项重复性实验证明了该方法的优势。