This paper studies the prediction of a target $\mathbf{z}$ from a pair of random variables $(\mathbf{x},\mathbf{y})$, where the ground-truth predictor is additive $\mathbb{E}[\mathbf{z} \mid \mathbf{x},\mathbf{y}] = f_\star(\mathbf{x}) +g_{\star}(\mathbf{y})$. We study the performance of empirical risk minimization (ERM) over functions $f+g$, $f \in F$ and $g \in G$, fit on a given training distribution, but evaluated on a test distribution which exhibits covariate shift. We show that, when the class $F$ is "simpler" than $G$ (measured, e.g., in terms of its metric entropy), our predictor is more resilient to heterogenous covariate shifts} in which the shift in $\mathbf{x}$ is much greater than that in $\mathbf{y}$. Our analysis proceeds by demonstrating that ERM behaves qualitatively similarly to orthogonal machine learning: the rate at which ERM recovers the $f$-component of the predictor has only a lower-order dependence on the complexity of the class $G$, adjusted for partial non-indentifiability introduced by the additive structure. These results rely on a novel H\"older style inequality for the Dudley integral which may be of independent interest. Moreover, we corroborate our theoretical findings with experiments demonstrating improved resilience to shifts in "simpler" features across numerous domains.
翻译:本文研究从一对随机变量 $(\mathbf{x},\mathbf{y})$ 预测目标 $\mathbf{z}$ 的问题,其中真实预测器为加性结构 $\mathbb{E}[\mathbf{z} \mid \mathbf{x},\mathbf{y}] = f_\star(\mathbf{x}) + g_{\star}(\mathbf{y})$。我们研究了经验风险最小化(ERM)在函数类 $f+g$(其中 $f \in F$,$g \in G$)上的性能,模型基于给定训练分布拟合,但在存在协变量偏移的测试分布上进行评估。我们证明,当 $F$ 类的复杂度(例如以度量熵衡量)低于 $G$ 类时,我们的预测器能够更好地抵御 $\mathbf{x}$ 的偏移远大于 $\mathbf{y}$ 的偏移的异质协变量偏移。我们的分析表明,ERM 的行为在定性上与正交机器学习类似:ERM 恢复预测器中 $f$ 分量的速率仅依赖于 $G$ 类复杂度的低阶项,并考虑了加性结构引入的部分非可识别性。这些结果依赖于一个新颖的、针对 Dudley 积分的 Hölder 型不等式,该不等式可能独立具有研究价值。此外,我们通过实验验证了理论发现,展示了在多个领域中针对“更简单”特征的偏移具有更强的鲁棒性。