Completing low-rank matrices from subsampled measurements has received much attention in the past decade. Existing works indicate that $\mathcal{O}(nr\log^2(n))$ datums are required to theoretically secure the completion of an $n \times n$ noisy matrix of rank $r$ with high probability, under some quite restrictive assumptions: (1) the underlying matrix must be incoherent; (2) observations follow the uniform distribution. The restrictiveness is partially due to ignoring the roles of the leverage score and the oracle information of each element. In this paper, we employ the leverage scores to characterize the importance of each element and significantly relax assumptions to: (1) not any other structure assumptions are imposed on the underlying low-rank matrix; (2) elements being observed are appropriately dependent on their importance via the leverage score. Under these assumptions, instead of uniform sampling, we devise an ununiform/biased sampling procedure that can reveal the ``importance'' of each observed element. Our proofs are supported by a novel approach that phrases sufficient optimality conditions based on the Golfing Scheme, which would be of independent interest to the wider areas. Theoretical findings show that we can provably recover an unknown $n\times n$ matrix of rank $r$ from just about $\mathcal{O}(nr\log^2 (n))$ entries, even when the observed entries are corrupted with a small amount of noisy information. The empirical results align precisely with our theories.
翻译:在过去的十年中,从子采样测量中完成低秩矩阵受到了广泛关注。现有研究表明,在以下相当严格的假设下,理论上需要大约 $\mathcal{O}(nr\log^2(n))$ 个数据点才能以高概率保证完成一个秩为 $r$ 的 $n \times n$ 含噪矩阵:(1)底层矩阵必须是非相干的;(2)观测值服从均匀分布。这种严格性部分源于忽略了杠杆分数和每个元素的神谕信息的作用。在本文中,我们利用杠杆分数来表征每个元素的重要性,并显著放宽假设至:(1)对底层低秩矩阵不施加任何其他结构假设;(2)被观测的元素通过杠杆分数适当地依赖于其重要性。在这些假设下,我们设计了一种非均匀/有偏的采样程序,可以揭示每个观测元素的“重要性”,而非均匀采样。我们的证明基于一种新颖的方法,该方法通过高尔夫方案(Golfing Scheme)表述充分最优性条件,这可能对更广泛的领域具有独立意义。理论结果表明,即使观测条目被少量噪声信息污染,我们也能从大约 $\mathcal{O}(nr\log^2(n))$ 个条目中可证明地恢复一个秩为 $r$ 的未知 $n \times n$ 矩阵。实证结果与我们的理论精确吻合。