We study robust regression under a contamination model in which covariates are clean while the responses may be corrupted in an adaptive manner. Unlike the classical Huber's contamination model, where both covariates and responses may be contaminated and consistent estimation is impossible when the contamination proportion is a non-vanishing constant, it turns out that the clean-covariate setting admits strictly improved statistical guarantees. Specifically, we show that the additional information in the clean covariates can be carefully exploited to construct an estimator that achieves a better estimation rate than that attainable under Huber contamination. In contrast to the Huber model, this improved rate implies consistency even when the contamination is a constant. A matching minimax lower bound is established using Fano's inequality together with the construction of contamination processes that match $m> 2$ distributions simultaneously, extending the previous two-point lower bound argument in Huber's setting. Despite the improvement over the Huber model from an information-theoretic perspective, we provide formal evidence -- in the form of Statistical Query and Low-Degree Polynomial lower bounds -- that the problem exhibits strong information-computation gaps. Our results strongly suggest that the information-theoretic improvements cannot be achieved by polynomial-time algorithms, revealing a fundamental gap between information-theoretic and computational limits in robust regression with clean covariates.
翻译:我们研究在一种污染模型下的稳健回归,其中协变量是干净的,而响应可能以自适应方式被污染。与经典的 Huber 污染模型(其中协变量和响应均可能被污染且当污染比例为非消失常数时无法实现一致估计)不同,干净协变量设定在统计上可提供严格改进的保证。具体而言,我们表明可以精心利用干净协变量中的额外信息来构建一个估计量,使其达到比 Huber 污染下更优的估计速率。与 Huber 模型相比,即使在污染为常数时,该改进速率也意味着一致性。利用 Fano 不等式以及构造同时匹配 $m>2$ 个分布的污染过程,我们建立了匹配的极小化极大下界,这扩展了 Huber 设定中先前的两点下界论证。尽管从信息论角度看较 Huber 模型有所改进,但我们以统计查询和低度多项式下界的形式提供了形式化证据,表明该问题存在强烈的信息-计算差距。我们的结果强烈表明,信息论上的改进无法通过多项式时间算法实现,揭示了在干净协变量稳健回归中信息论极限与计算极限之间的根本差距。