Motivated by a recent literature on the double-descent phenomenon in machine learning, we consider highly over-parameterized models in causal inference, including synthetic control with many control units. In such models, there may be so many free parameters that the model fits the training data perfectly. We first investigate high-dimensional linear regression for imputing wage data and estimating average treatment effects, where we find that models with many more covariates than sample size can outperform simple ones. We then document the performance of high-dimensional synthetic control estimators with many control units. We find that adding control units can help improve imputation performance even beyond the point where the pre-treatment fit is perfect. We provide a unified theoretical perspective on the performance of these high-dimensional models. Specifically, we show that more complex models can be interpreted as model-averaging estimators over simpler ones, which we link to an improvement in average performance. This perspective yields concrete insights into the use of synthetic control when control units are many relative to the number of pre-treatment periods.
翻译:受机器学习中双下降现象的近期文献启发,我们研究了因果推断中高度过参数化的模型,包括包含大量对照单元的合成控制方法。在这类模型中,自由参数可能过多,以至于模型完美拟合训练数据。我们首先探讨了用于工资数据插补和平均处理效应估计的高维线性回归,发现当协变量数量远大于样本量时,这类模型的表现可优于简单模型。随后我们记录了含大量对照单元的高维合成控制估计量的性能,发现即使在预处理拟合达到完美后,增加对照单元仍能改善插补效果。我们为这些高维模型的性能提供了统一的理论视角:具体而言,我们证明更复杂的模型可被解释为对简单模型的模型平均估计量,这与其平均性能的提升相关联。该视角为处理对照单元数量远多于预处理期数的合成控制方法提供了具体洞见。