High-dimensional imbalanced data poses a machine learning challenge. In the absence of sufficient or high-quality labels, unsupervised feature selection methods are crucial for the success of subsequent algorithms. Therefore, we introduce a Marginal Laplacian Score (MLS), a modification of the well known Laplacian Score (LS) tailored to better address imbalanced data. We introduce an assumption that the minority class or anomalous appear more frequently in the margin of the features. Consequently, MLS aims to preserve the local structure of the dataset's margin. We propose its integration into modern feature selection methods that utilize the Laplacian score. We integrate the MLS algorithm into the Differentiable Unsupervised Feature Selection (DUFS), resulting in DUFS-MLS. The proposed methods demonstrate robust and improved performance on synthetic and public datasets.
翻译:高维不平衡数据对机器学习构成挑战。在缺乏足够或高质量标签的情况下,无监督特征选择方法对后续算法的成功至关重要。为此,我们提出边际拉普拉斯得分(MLS),这是对经典拉普拉斯得分(LS)的改进,旨在更好地处理不平衡数据。我们引入一个假设:少数类或异常值更频繁地出现在特征的边际区域。因此,MLS旨在保留数据集边际的局部结构。我们提出将其集成到利用拉普拉斯得分的现代特征选择方法中。我们将MLS算法集成到可微无监督特征选择(DUFS)中,形成DUFS-MLS。所提出的方法在合成数据集和公开数据集上展现出稳健且更优的性能。