This paper is concerned with estimating the column subspace of a low-rank matrix $\boldsymbol{X}^\star \in \mathbb{R}^{n_1\times n_2}$ from contaminated data. How to obtain optimal statistical accuracy while accommodating the widest range of signal-to-noise ratios (SNRs) becomes particularly challenging in the presence of heteroskedastic noise and unbalanced dimensionality (i.e., $n_2\gg n_1$). While the state-of-the-art algorithm $\textsf{HeteroPCA}$ emerges as a powerful solution for solving this problem, it suffers from "the curse of ill-conditioning," namely, its performance degrades as the condition number of $\boldsymbol{X}^\star$ grows. In order to overcome this critical issue without compromising the range of allowable SNRs, we propose a novel algorithm, called $\textsf{Deflated-HeteroPCA}$, that achieves near-optimal and condition-number-free theoretical guarantees in terms of both $\ell_2$ and $\ell_{2,\infty}$ statistical accuracy. The proposed algorithm divides the spectrum of $\boldsymbol{X}^\star$ into well-conditioned and mutually well-separated subblocks, and applies $\textsf{HeteroPCA}$ to conquer each subblock successively. Further, an application of our algorithm and theory to two canonical examples -- the factor model and tensor PCA -- leads to remarkable improvement for each application.
翻译:本文研究从含噪数据中估计低秩矩阵 $\boldsymbol{X}^\star \in \mathbb{R}^{n_1\times n_2}$ 的列子空间问题。在异方差噪声和维度不平衡(即 $n_2\gg n_1$)的情况下,如何在兼顾最广泛信噪比(SNR)范围的同时获得最优统计精度,变得尤为具有挑战性。尽管当前最先进的算法 $\textsf{HeteroPCA}$ 为解决该问题提供了强大方案,但其受困于“病态性诅咒”——即性能随 $\boldsymbol{X}^\star$ 条件数的增大而显著下降。为克服这一关键问题且不牺牲可允许的SNR范围,本文提出一种名为 $\textsf{Deflated-HeteroPCA}$ 的新算法,该算法在 $\ell_2$ 和 $\ell_{2,\infty}$ 统计精度上均实现了近最优且无条件数依赖的理论保证。所提算法将 $\boldsymbol{X}^\star$ 的谱分解为良态且彼此充分分离的子块,并依次应用 $\textsf{HeteroPCA}$ 攻克每个子块。进一步地,将其算法与理论应用于两个经典范例——因子模型和张量 PCA——在各自应用中均取得了显著性能提升。