It is well-known that the statistical performance of Lasso can suffer significantly when the covariates of interest have strong correlations. In particular, the prediction error of Lasso becomes much worse than computationally inefficient alternatives like Best Subset Selection. Due to a large conjectured computational-statistical tradeoff in the problem of sparse linear regression, it may be impossible to close this gap in general. In this work, we propose a natural sparse linear regression setting where strong correlations between covariates arise from unobserved latent variables. In this setting, we analyze the problem caused by strong correlations and design a surprisingly simple fix. While Lasso with standard normalization of covariates fails, there exists a heterogeneous scaling of the covariates with which Lasso will suddenly obtain strong provable guarantees for estimation. Moreover, we design a simple, efficient procedure for computing such a "smart scaling." The sample complexity of the resulting "rescaled Lasso" algorithm incurs (in the worst case) quadratic dependence on the sparsity of the underlying signal. While this dependence is not information-theoretically necessary, we give evidence that it is optimal among the class of polynomial-time algorithms, via the method of low-degree polynomials. This argument reveals a new connection between sparse linear regression and a special version of sparse PCA with a near-critical negative spike. The latter problem can be thought of as a real-valued analogue of learning a sparse parity. Using it, we also establish the first computational-statistical gap for the closely related problem of learning a Gaussian Graphical Model.
翻译:众所周知,当关注的协变量之间存在强相关性时,Lasso的统计性能会显著下降。特别地,Lasso的预测误差远不如计算效率较低的最优子集选择等替代方法。由于稀疏线性回归问题中普遍存在的计算统计权衡假设,这一差距在一般情况下可能无法弥补。本文提出了一种自然的稀疏线性回归场景,其中协变量之间的强相关性源于未观测到的潜变量。在此场景中,我们分析了强相关性引发的问题,并设计了一个出人意料的简单修正方案。虽然采用标准归一化协变量的Lasso方法失效,但存在一种异质性的协变量缩放方式,使得Lasso能够突然获得强大的可证明估计保证。此外,我们设计了一种简单高效的计算流程来实现这种"智能缩放"。得到的"重缩放Lasso"算法的样本复杂度(在最坏情况下)与底层信号的稀疏度呈二次关系。虽然这种依赖关系在信息论上并非必要,但通过低阶多项式方法,我们证明了其在多项式时间算法类中的最优性。这一论证揭示了稀疏线性回归与具有近临界负尖峰的稀疏主成分分析特殊版本之间的新联系。后者可视为学习稀疏奇偶性的实值类比。借此,我们还首次建立了学习高斯图模型这一密切相关问题的计算统计差距。