In this manuscript we derive the optimal out-of-sample causal predictor for a linear system that has been observed in $k+1$ within-sample environments. In this model we consider $k$ shifted environments and one observational environment. Each environment corresponds to a linear structural equation model (SEM) with its own shift and noise vector, both in $L^2$. The strength of the shifts can be put in a certain order, and we may therefore speak of all shifts that are less or equally strong than a given shift. We consider the space of all shifts are $\gamma$ times less or equally strong than any weighted average of the observed shift vectors with weights on the unit sphere. For each $\beta\in\mathbb{R}^p$ we show that the supremum of the risk functions $R_{\tilde{A}}(\beta)$ over $\tilde{A}\in C^\gamma$ has a worst-risk decomposition into a (positive) linear combination of risk functions, depending on $\gamma$. We then define the causal regularizer, $\beta_\gamma$, as the argument $\beta$ that minimizes this risk. The main result of the paper is that this regularizer can be consistently estimated with a plug-in estimator outside a set of zero Lebesgue measure in the parameter space. A practical obstacle for such estimation is that it involves the solution of a general degree polynomial which cannot be done explicitly. Therefore we also prove that an approximate plug-in estimator using the bisection method is also consistent. An interesting by-product of the proof of the main result is that the plug-in estimation of the argmin of the maxima of a finite set of quadratic risk functions is consistent outside a set of zero Lebesgue measure in the parameter space.
翻译:本文推导了在 $k+1$ 个样本内环境中观测到的线性系统的最优样本外因果预测器。在该模型中,我们考虑 $k$ 个偏移环境和一个观测环境。每个环境对应一个线性结构方程模型(SEM),其具有各自的偏移向量和噪声向量(两者均属于 $L^2$ 空间)。这些偏移的强度可按一定顺序排列,因此我们可以讨论所有弱于或等于给定偏移的偏移。我们考虑所有偏移的集合,这些偏移的强度不超过观测偏移向量在单位球面上任意加权平均的 $\gamma$ 倍。对于每个 $\beta\in\mathbb{R}^p$,我们证明风险函数 $R_{\tilde{A}}(\beta)$ 在 $\tilde{A}\in C^\gamma$ 上的上确界具有一种最坏情况风险分解形式,即表示为风险函数的(正)线性组合,且该组合依赖于 $\gamma$。随后,我们将因果正则化器 $\beta_\gamma$ 定义为最小化该风险的参数 $\beta$。本文的主要结果是:在参数空间中除去一个勒贝格测度为零的集合外,该正则化器可通过代入法估计器得到一致估计。此类估计的实际障碍在于其涉及一个一般次数多项式的求解,而该求解无法显式完成。因此,我们还证明了一种使用二分法的近似代入法估计器也是一致估计。主要结果证明的一个有趣副产品是:在参数空间中除去一个零勒贝格测度集后,有限个二次风险函数最大值的最小值点的代入法估计具有一致性。