This paper presents a new filter method for unsupervised feature selection. This method is particularly effective on imbalanced multi-class dataset, as in case of clusters of different anomaly types. Existing methods usually involve the variance of the features, which is not suitable when the different types of observations are not represented equally. Our method, based on Spearman's Rank Correlation between distances on the observations and on feature values, avoids this drawback. The performance of the method is measured on several clustering problems and is compared with existing filter methods suitable for unsupervised data.
翻译:本文提出一种新的非监督特征选择过滤方法。该方法对不平衡多类数据集(如包含多种异常类型的聚类场景)尤为有效。现有方法通常依赖于特征方差,但当观测类型分布不均时,该指标并不适用。本文方法基于观测距离与特征值之间的斯皮尔曼等级相关(Spearman's Rank Correlation),有效避免了上述缺陷。通过在多个聚类问题上的性能评估,并与现有适用于非监督数据的过滤方法进行对比,验证了该方法的有效性。