Manifold learning approaches seek the intrinsic, low-dimensional data structure within a high-dimensional space. Mainstream manifold learning algorithms, such as Isomap, UMAP, $t$-SNE, Diffusion Map, and Laplacian Eigenmaps do not use data labels and are thus considered unsupervised. Existing supervised extensions of these methods are limited to classification problems and fall short of uncovering meaningful embeddings due to their construction using order non-preserving, class-conditional distances. In this paper, we show the weaknesses of class-conditional manifold learning quantitatively and visually and propose an alternate choice of kernel for supervised dimensionality reduction using a data-geometry-preserving variant of random forest proximities as an initialization for manifold learning methods. We show that local structure preservation using these proximities is near universal across manifold learning approaches and global structure is properly maintained using diffusion-based algorithms.
翻译:流形学习方法旨在揭示高维空间中内在的低维数据结构。主流的流形学习算法,如Isomap、UMAP、$t$-SNE、扩散映射和拉普拉斯特征映射,由于不利用数据标签,因此被视为无监督方法。这些方法现有的监督扩展仅限于分类问题,并且由于使用非保序的类条件距离,未能揭示有意义的嵌入。本文从定量和视觉角度展示了类条件流形学习的局限性,并提出一种监督降维的替代核函数选择,即采用保数据几何的随机森林邻近性变体作为流形学习方法的初始化。研究表明,基于这些邻近性的局部结构保留在流形学习方法中近乎普遍适用,而全局结构则通过基于扩散的算法得以恰当维持。