Spectral clustering and diffusion maps are celebrated dimensionality reduction algorithms built on eigen-elements related to the diffusive structure of the data. The core of these procedures is the approximation of a Laplacian through a graph kernel approach, however this local average construction is known to be cursed by the high-dimension d. In this article, we build a different estimator of the Laplacian, via a reproducing kernel Hilbert space method, which adapts naturally to the regularity of the problem. We provide non-asymptotic statistical rates proving that the kernel estimator we build can circumvent the curse of dimensionality. Finally we discuss techniques (Nystr\"om subsampling, Fourier features) that enable to reduce the computational cost of the estimator while not degrading its overall performance.
翻译:谱聚类和扩散映射是基于与数据扩散结构相关的特征元素构建的经典降维算法。这些方法的核心是通过图核方法近似拉普拉斯算子,然而这种局部平均构造在高维空间d中已知会受到维度灾难的影响。本文通过再生核希尔伯特空间方法,构建了一种不同的拉普拉斯算子估计量,该估计量自然地适应问题的正则性。我们提供了非渐近统计速率,证明所构建的核估计量能够规避维度灾难。最后,我们讨论了能够降低估计量计算成本且不损害整体性能的技术(Nyström子采样、傅里叶特征)。