We investigate the high-dimensional properties of robust regression estimators in the presence of heavy-tailed contamination of both the covariates and response functions. In particular, we provide a sharp asymptotic characterisation of M-estimators trained on a family of elliptical covariate and noise data distributions including cases where second and higher moments do not exist. We show that, despite being consistent, the Huber loss with optimally tuned location parameter $\delta$ is suboptimal in the high-dimensional regime in the presence of heavy-tailed noise, highlighting the necessity of further regularisation to achieve optimal performance. This result also uncovers the existence of a curious transition in $\delta$ as a function of the sample complexity and contamination. Moreover, we derive the decay rates for the excess risk of ridge regression. We show that, while it is both optimal and universal for noise distributions with finite second moment, its decay rate can be considerably faster when the covariates' second moment does not exist. Finally, we show that our formulas readily generalise to a richer family of models and data distributions, such as generalised linear estimation with arbitrary convex regularisation trained on mixture models.
翻译:我们研究了在协变量和响应函数均存在重尾污染时,稳健回归估计器的高维性质。具体而言,我们给出了M估计器在椭圆协变量和噪声数据分布族(包括二阶及更高阶矩不存在的情形)上的精确渐近刻画。研究表明,尽管Huber损失结合最优调节位置参数$\delta$具有一致性,但在高维机制下重尾噪声中其表现次优,这凸显了实现最优性能需要进一步正则化的必要性。该结果还揭示了$\delta$随样本复杂度和污染程度变化的特殊转变现象。此外,我们推导了岭回归过量风险的衰减速率。结果表明,尽管对于具有有限二阶矩的噪声分布而言该衰减速率兼具最优性与普适性,但当协变量二阶矩不存在时,其衰减速率可能显著加快。最后,我们证明了公式可直接推广至更丰富的模型与数据分布族,例如基于混合模型训练的任意凸正则化广义线性估计。