We give a new framework for solving the fundamental problem of low-rank matrix completion, i.e., approximating a rank-$r$ matrix $\mathbf{M} \in \mathbb{R}^{m \times n}$ (where $m \ge n$) from random observations. First, we provide an algorithm which completes $\mathbf{M}$ on $99\%$ of rows and columns under no further assumptions on $\mathbf{M}$ from $\approx mr$ samples and using $\approx mr^2$ time. Then, assuming the row and column spans of $\mathbf{M}$ satisfy additional regularity properties, we show how to boost this partial completion guarantee to a full matrix completion algorithm by aggregating solutions to regression problems involving the observations. In the well-studied setting where $\mathbf{M}$ has incoherent row and column spans, our algorithms complete $\mathbf{M}$ to high precision from $mr^{2+o(1)}$ observations in $mr^{3 + o(1)}$ time (omitting logarithmic factors in problem parameters), improving upon the prior state-of-the-art [JN15] which used $\approx mr^5$ samples and $\approx mr^7$ time. Under an assumption on the row and column spans of $\mathbf{M}$ we introduce (which is satisfied by random subspaces with high probability), our sample complexity improves to an almost information-theoretically optimal $mr^{1 + o(1)}$, and our runtime improves to $mr^{2 + o(1)}$. Our runtimes have the appealing property of matching the best known runtime to verify that a rank-$r$ decomposition $\mathbf{U}\mathbf{V}^\top$ agrees with the sampled observations. We also provide robust variants of our algorithms that, given random observations from $\mathbf{M} + \mathbf{N}$ with $\|\mathbf{N}\|_{F} \le \Delta$, complete $\mathbf{M}$ to Frobenius norm distance $\approx r^{1.5}\Delta$ in the same runtimes as the noiseless setting. Prior noisy matrix completion algorithms [CP10] only guaranteed a distance of $\approx \sqrt{n}\Delta$.
翻译:我们提出一个新框架,用于解决低秩矩阵补全这一基本问题,即从随机观测中逼近秩为$r$的矩阵$\mathbf{M} \in \mathbb{R}^{m \times n}$(其中$m \ge n$)。首先,我们给出一个算法,它能在无需对$\mathbf{M}$附加任何假设的情况下,从约$mr$个样本中、用时约$mr^2$完成$99\%$的行和列上的$\mathbf{M}$补全。接着,假设$\mathbf{M}$的行空间和列空间满足额外的正则性质,我们展示如何通过聚合涉及观测的回归问题的解,将这种部分补全保证提升为完整的矩阵补全算法。在广为研究的设置中,即$\mathbf{M}$具有非相干的行空间和列空间时,我们的算法能从$mr^{2+o(1)}$个观测中、在$mr^{3 + o(1)}$时间内(忽略问题参数的对数因子)高精度地完成$\mathbf{M}$的补全,这改进了先前的最先进方法[JN15],后者使用了约$mr^5$个样本和约$mr^7$时间。在我们引入的对$\mathbf{M}$的行空间和列空间的假设下(该假设以高概率被随机子空间满足),我们的样本复杂度提升到接近信息论最优的$mr^{1 + o(1)}$,运行时间提升到$mr^{2 + o(1)}$。我们的运行时间具有与验证一个秩为$r$的分解$\mathbf{U}\mathbf{V}^\top$是否与采样观测一致的已知最优运行时间相匹配的吸引特性。我们还提供算法的鲁棒变体,给定来自$\mathbf{M} + \mathbf{N}$(其中$\|\mathbf{N}\|_{F} \le \Delta$)的随机观测时,这些变体能在与无噪声设置相同的运行时间内,以Frobenius范数距离约$r^{1.5}\Delta$完成$\mathbf{M}$的补全。先前的有噪矩阵补全算法[CP10]仅能保证约$\sqrt{n}\Delta$的距离。