The Chebyshev or $\ell_{\infty}$ estimator is an unconventional alternative to the ordinary least squares in solving linear regressions. It is defined as the minimizer of the $\ell_{\infty}$ objective function \begin{align*} \hat{\boldsymbol{\beta}} := \arg\min_{\boldsymbol{\beta}} \|\boldsymbol{Y} - \mathbf{X}\boldsymbol{\beta}\|_{\infty}. \end{align*} The asymptotic distribution of the Chebyshev estimator under fixed number of covariates was recently studied (Knight, 2020), yet finite sample guarantees and generalizations to high-dimensional settings remain open. In this paper, we develop non-asymptotic upper bounds on the estimation error $\|\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}^*\|_2$ for a Chebyshev estimator $\hat{\boldsymbol{\beta}}$, in a regression setting with uniformly distributed noise $\varepsilon_i\sim U([-a,a])$ where $a$ is either known or unknown. With relatively mild assumptions on the (random) design matrix $\mathbf{X}$, we can bound the error rate by $\frac{C_p}{n}$ with high probability, for some constant $C_p$ depending on the dimension $p$ and the law of the design. Furthermore, we illustrate that there exist designs for which the Chebyshev estimator is (nearly) minimax optimal. On the other hand we also argue that there exist designs for which this estimator behaves sub-optimally in terms of the constant $C_p$'s dependence on $p$. In addition we show that "Chebyshev's LASSO" has advantages over the regular LASSO in high dimensional situations, provided that the noise is uniform. Specifically, we argue that it achieves a much faster rate of estimation under certain assumptions on the growth rate of the sparsity level and the ambient dimension with respect to the sample size.
翻译:Chebyshev估计量(即$\ell_{\infty}$估计量)是求解线性回归时替代普通最小二乘法的非传统方法,其定义为最小化$\ell_{\infty}$目标函数:
\begin{align*}
\hat{\boldsymbol{\beta}} := \arg\min_{\boldsymbol{\beta}} \|\boldsymbol{Y} - \mathbf{X}\boldsymbol{\beta}\|_{\infty}.
\end{align*}
近期研究(Knight, 2020)虽已探索了协变量数量固定时Chebyshev估计量的渐近分布,但其有限样本保证及对高维场景的推广仍属开放问题。本文针对具有均匀分布噪声$\varepsilon_i\sim U([-a,a])$(其中$a$已知或未知)的回归场景,建立了Chebyshev估计量$\hat{\boldsymbol{\beta}}$的估计误差$\|\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}^*\|_2$的非渐近上界。在对(随机)设计矩阵$\mathbf{X}$施加相对温和的假设条件下,我们以高概率将误差率界定为$\frac{C_p}{n}$,其中常数$C_p$依赖于维度$p$及设计律。此外,我们论证了存在使Chebyshev估计量达到(接近)极小极大最优的设计,同时也存在该估计量因常数$C_p$对$p$的依赖性而表现次优的设计。进一步,我们证明均匀噪声下"Chebyshev LASSO"在高维场景中优于常规LASSO:具体而言,在关于稀疏度水平与环境维度相对于样本量增长率的特定假设下,其可实现更快的估计速率。