Principal component analysis (PCA) is a simple and popular tool for processing high-dimensional data. We investigate its effectiveness for matrix denoising. We assume i.i.d. high dimensional Gaussian noises with standard deviation $\sigma$ are added to clean data generated from a low dimensional subspace. We show that the distance between each pair of PCA-denoised data point and the clean data point is uniformly bounded by $\Otilde(\sigma)$, assuming a low-rank data matrix with mild singular value assumptions. We show such a condition could arise even if the data lies on curves. We then provide a general lower bound for the error of the denoised data matrix, which indicates PCA denoising gives a uniform error bound that is rate-optimal. Furthermore, we examine how the error bound impacts downstream applications such as empirical risk minimization, clustering, and manifold learning. Numerical results validate our theoretical findings and reveal the importance of the uniform error.
翻译:主成分分析(PCA)是一种简单且流行的处理高维数据的工具。我们研究其在矩阵去噪中的有效性。我们假设添加了标准差为$\sigma$的独立同分布高维高斯噪声到由低维子空间生成的干净数据中。我们证明,在低秩数据矩阵具有温和奇异值假设的条件下,每对PCA去噪数据点与干净数据点之间的距离被$\Otilde(\sigma)$一致有界。我们表明,即使数据位于曲线上,这种条件也可能成立。随后,我们给出了去噪数据矩阵误差的通用下界,表明PCA去噪提供了速率最优的一致误差界。此外,我们研究了误差界对下游应用(如经验风险最小化、聚类和流形学习)的影响。数值结果验证了我们的理论发现,并揭示了均匀误差的重要性。