In recent decades, a number of ways of dealing with causality in practice, such as propensity score matching, the PC algorithm and invariant causal prediction, have been introduced. Besides its interpretational appeal, the causal model provides the best out-of-sample prediction guarantees. In this paper, we study the identification of causal-like models from in-sample data that provide out-of-sample risk guarantees when predicting a target variable from a set of covariates. Whereas ordinary least squares provides the best in-sample risk with limited out-of-sample guarantees, causal models have the best out-of-sample guarantees but achieve an inferior in-sample risk. By defining a trade-off of these properties, we introduce $\textit{causal regularization}$. As the regularization is increased, it provides estimators whose risk is more stable across sub-samples at the cost of increasing their overall in-sample risk. The increased risk stability is shown to lead to out-of-sample risk guarantees. We provide finite sample risk bounds for all models and prove the adequacy of cross-validation for attaining these bounds.
翻译:近几十年来,多种在实践中处理因果性的方法被提出,例如倾向得分匹配、PC算法以及不变因果预测。因果模型除了具有解释性方面的吸引力,还提供了最佳的样本外预测保证。本文研究如何从样本内数据识别类因果模型,这些模型在基于一组协变量预测目标变量时能够提供样本外风险保证。普通最小二乘法在样本内风险方面表现最佳,但样本外保证有限;而因果模型具有最佳的样本外保证,但样本内风险较差。通过定义这些性质的权衡,我们引入了\textit{因果正则化}。随着正则化强度的增加,其估计量在不同子样本间的风险更稳定,但代价是整体样本内风险上升。研究表明,这种增强的风险稳定性可带来样本外风险保证。我们为所有模型提供了有限样本风险界,并证明了交叉验证在达到这些界时的充分性。