Principal component analysis (PCA) is a widely employed statistical tool used primarily for dimensionality reduction. However, it is known to be adversely affected by the presence of outlying observations in the sample, which is quite common. Robust PCA methods using M-estimators have theoretical benefits, but their robustness drop substantially for high dimensional data. On the other end of the spectrum, robust PCA algorithms solving principal component pursuit or similar optimization problems have high breakdown, but lack theoretical richness and demand high computational power compared to the M-estimators. We introduce a novel robust PCA estimator based on the minimum density power divergence estimator. This combines the theoretical strength of the M-estimators and the minimum divergence estimators with a high breakdown guarantee regardless of data dimension. We present a computationally efficient algorithm for this estimate. Our theoretical findings are supported by extensive simulations and comparisons with existing robust PCA methods. We also showcase the proposed algorithm's applicability on two benchmark datasets and a credit card transactions dataset for fraud detection.
翻译:主成分分析(PCA)是一种广泛使用的统计工具,主要用于降维。然而,它容易受到样本中存在异常值(这相当常见)的不利影响。使用M估计的鲁棒PCA方法具有理论优势,但其稳健性在高维数据中会大幅下降。另一方面,解决主成分追踪或类似优化问题的鲁棒PCA算法具有高崩溃点,但与M估计相比,缺乏理论丰富性且计算需求较高。我们提出了一种基于最小密度功率散度估计的新型鲁棒PCA估计器。这结合了M估计和最小散度估计的理论优势,同时无论数据维度如何,都能提供高崩溃点保证。我们提出了一种计算高效的算法来实现这一估计。我们的理论发现得到了大量模拟实验和与现有鲁棒PCA方法比较的支持。我们还展示了所提出算法在两个基准数据集和一个用于欺诈检测的信用卡交易数据集上的适用性。