An increasing number of data science and machine learning problems rely on computation with tensors, which better capture the multi-way relationships and interactions of data than matrices. When tapping into this critical advantage, a key challenge is to develop computationally efficient and provably correct algorithms for extracting useful information from tensor data that are simultaneously robust to corruptions and ill-conditioning. This paper tackles tensor robust principal component analysis (RPCA), which aims to recover a low-rank tensor from its observations contaminated by sparse corruptions, under the Tucker decomposition. To minimize the computation and memory footprints, we propose to directly recover the low-dimensional tensor factors -- starting from a tailored spectral initialization -- via scaled gradient descent (ScaledGD), coupled with an iteration-varying thresholding operation to adaptively remove the impact of corruptions. Theoretically, we establish that the proposed algorithm converges linearly to the true low-rank tensor at a constant rate that is independent with its condition number, as long as the level of corruptions is not too large. Empirically, we demonstrate that the proposed algorithm achieves better and more scalable performance than state-of-the-art matrix and tensor RPCA algorithms through synthetic experiments and real-world applications.
翻译:日益增多的数据科学和机器学习问题依赖于张量计算,相比矩阵,张量能更好地捕捉数据的多路关系与交互。在利用这一关键优势时,一个核心挑战是开发计算高效且可证明正确的算法,用于从受污染和病态条件影响的张量数据中提取有用信息。本文研究在Tucker分解框架下,旨在从被稀疏污染观测中恢复低秩张量的张量稳健主成分分析。为最小化计算和内存开销,我们提出直接恢复低维张量因子——从定制的谱初始化开始——通过缩放梯度下降法,并耦合随迭代次数变化的阈值操作以自适应消除污染的影响。理论上,我们证明该算法以与条件数无关的恒定速率线性收敛至真实低秩张量,前提是污染程度不太大。实验上,通过合成实验和实际应用,我们证明该算法在性能与可扩展性上均优于当前最先进的矩阵与张量稳健主成分分析算法。