A CUR factorization is often utilized as a substitute for the singular value decomposition (SVD), especially when a concrete interpretation of the singular vectors is challenging. Moreover, if the original data matrix possesses properties like nonnegativity and sparsity, a CUR decomposition can better preserve them compared to the SVD. An essential aspect of this approach is the methodology used for selecting a subset of columns and rows from the original matrix. This study investigates the effectiveness of \emph{one-round sampling} and iterative subselection techniques and introduces new iterative subselection strategies based on iterative SVDs. One provably appropriate technique for index selection in constructing a CUR factorization is the discrete empirical interpolation method (DEIM). Our contribution aims to improve the approximation quality of the DEIM scheme by iteratively invoking it in several rounds, in the sense that we select subsequent columns and rows based on the previously selected ones. Thus, we modify $A$ after each iteration by removing the information that has been captured by the previously selected columns and rows. We also discuss how iterative procedures for computing a few singular vectors of large data matrices can be integrated with the new iterative subselection strategies. We present the results of numerical experiments, providing a comparison of one-round sampling and iterative subselection techniques, and demonstrating the improved approximation quality associated with using the latter.
翻译:CUR分解通常被用作奇异值分解(SVD)的替代方法,尤其是在对奇异向量进行具体解释较为困难的情况下。此外,如果原始数据矩阵具有非负性和稀疏性等特性,CUR分解相比SVD能更好地保留这些特性。该方法的一个关键方面是用于从原始矩阵中选择列和行子集的技术。本研究探讨了“一轮采样”和迭代子选择技术的有效性,并引入了基于迭代SVD的新的迭代子选择策略。在构建CUR分解时,一种已被证明适用于索引选择的技术是离散经验插值法(DEIM)。我们的贡献旨在通过多轮迭代调用DEIM方案来提高其逼近质量,即根据先前选择的列和行来选取后续的列和行。因此,我们在每次迭代后通过移除先前选择的列和行所捕获的信息来修改矩阵A。我们还讨论了如何将对大型数据矩阵计算少量奇异向量的迭代过程与新的迭代子选择策略相结合。我们展示了数值实验的结果,对一轮采样和迭代子选择技术进行了比较,并证明了使用后者时逼近质量的改善。