Real-world datasets are often of high dimension and effected by the curse of dimensionality. This hinders their comprehensibility and interpretability. To reduce the complexity feature selection aims to identify features that are crucial to learn from said data. While measures of relevance and pairwise similarities are commonly used, the curse of dimensionality is rarely incorporated into the process of selecting features. Here we step in with a novel method that identifies the features that allow to discriminate data subsets of different sizes. By adapting recent work on computing intrinsic dimensionalities, our method is able to select the features that can discriminate data and thus weaken the curse of dimensionality. Our experiments show that our method is competitive and commonly outperforms established feature selection methods. Furthermore, we propose an approximation that allows our method to scale to datasets consisting of millions of data points. Our findings suggest that features that discriminate data and are connected to a low intrinsic dimensionality are meaningful for learning procedures.
翻译:现实世界的数据集通常具有高维特征,并受到维度灾难的影响,这阻碍了数据的可理解性与可解释性。为降低复杂性,特征选择旨在识别对从该数据中学习至关重要的特征。虽然相关性度量和成对相似性度量被广泛使用,但维度灾难却很少被纳入特征选择过程。为此,我们提出一种新方法,用于识别能够区分不同规模数据子集的特征。通过借鉴近期关于内在维度计算的研究,该方法能筛选出可有效区分数据的特征,从而削弱维度灾难的影响。实验表明,我们的方法具有竞争力,且通常优于传统特征选择方法。此外,我们提出一种近似计算方案,使该方法能够扩展到包含数百万个数据点的数据集。研究结果表明,能够区分数据且与低内在维度相关联的特征对学习过程具有实际意义。