Sparse principal component analysis (PCA) aims at mapping large dimensional data to a linear subspace of lower dimension. By imposing loading vectors to be sparse, it performs the double duty of dimension reduction and variable selection. Sparse PCA algorithms are usually expressed as a trade-off between explained variance and sparsity of the loading vectors (i.e., number of selected variables). As a high explained variance is not necessarily synonymous with relevant information, these methods are prone to select irrelevant variables. To overcome this issue, we propose an alternative formulation of sparse PCA driven by the false discovery rate (FDR). We then leverage the Terminating-Random Experiments (T-Rex) selector to automatically determine an FDR-controlled support of the loading vectors. A major advantage of the resulting T-Rex PCA is that no sparsity parameter tuning is required. Numerical experiments and a stock market data example demonstrate a significant performance improvement.
翻译:稀疏主成分分析旨在将高维数据映射到较低维度的线性子空间。通过强制载荷向量稀疏化,它同时实现降维与变量选择的双重功能。稀疏主成分分析算法通常被表述为解释方差与载荷向量稀疏度(即所选变量数量)之间的权衡。由于高解释方差未必等同于相关信息,此类方法易导致无关变量的误选。为解决这一问题,我们提出以错误发现率驱动的稀疏主成分分析替代方案。进而利用终止-随机实验选择器自动确定经FDR控制的载荷向量支撑集。由此产生的T-Rex PCA方法具有无需调节稀疏度参数的重要优势。数值实验与股票市场数据案例表明该方法性能显著提升。