We study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the performances of the empirical risk minimisation (ERM) using $\ell_2$-regularised $\ell_2$, $\ell_1$, and Huber losses, which are the standard approach to such problems. We focus on two metrics for the performance: the generalisation error to similar datasets with outliers, and the estimation error of the original, unpolluted function. Our results are compared with the information theoretic Bayes-optimal estimation bound. For the generalization error, we find that optimally-regularised ERM is asymptotically consistent in the large sample complexity limit if one perform a simple calibration, and compute the rates of convergence. For the estimation error however, we show that due to a norm calibration mismatch, the consistency of the estimator requires an oracle estimate of the optimal norm, or the presence of a cross-validation set not corrupted by the outliers. We examine in detail how performance depends on the loss function and on the degree of outlier corruption in the training set and identify a region of parameters where the optimal performance of the Huber loss is identical to that of the $\ell_2$ loss, offering insights into the use cases of different loss functions.
翻译:我们研究了高维情形下鲁棒线性回归问题,其中维度 $d$ 与数据点数量 $n$ 以固定比例 $\alpha=n/d$ 共同发散,并考察了包含离群值的数据模型。针对该问题的标准方法——使用 $\ell_2$ 正则化的 $\ell_2$ 损失、$\ell_1$ 损失以及 Huber 损失的经验风险最小化(ERM),我们给出了其性能的精确渐近分析。我们关注两种性能指标:对含离群值相似数据集的泛化误差,以及对原始无污染函数的估计误差。我们的结果与信息论最优贝叶斯估计界进行了比较。对于泛化误差,我们发现若进行简单校准,最优正则化的 ERM 在大样本复杂度极限下渐近一致,并给出了收敛速率。然而对于估计误差,我们表明由于范数校准失配,估计量的一致性需要预知最优范数的 oracle 估计,或存在未被离群值污染的交差验证集。我们详细考察了性能如何依赖于损失函数以及训练集中离群值污染程度,并识别出一个参数区域,其中 Huber 损失的最优性能与 $\ell_2$ 损失相同,从而为不同损失函数的使用场景提供了洞见。