The optimal number of clusters is one of the main concerns when applying cluster analysis. Several cluster validity indexes have been introduced to address this problem. However, in some situations, there is more than one option that can be chosen as the final number of clusters. This aspect has been overlooked by most of the existing works in this area. In this study, we introduce a correlation-based fuzzy cluster validity index known as the Wiroonsri-Preedasawakul (WP) index. This index is defined based on the correlation between the actual distance between a pair of data points and the distance between adjusted centroids with respect to that pair. We evaluate and compare the performance of our index with several existing indexes, including Xie-Beni, Pakhira-Bandyopadhyay-Maulik, Tang, Wu-Li, generalized C, and Kwon2. We conduct this evaluation on four types of datasets: artificial datasets, real-world datasets, simulated datasets with ranks, and image datasets, using the fuzzy c-means algorithm. Overall, the WP index outperforms most, if not all, of these indexes in terms of accurately detecting the optimal number of clusters and providing accurate secondary options. Moreover, our index remains effective even when the fuzziness parameter $m$ is set to a large value. Our R package called WPfuzzyCVIs used in this work is also available in https://github.com/nwiroonsri/WPfuzzyCVIs.
翻译:聚类分析中的一个主要关注点是确定最优聚类数。为此,已有多种聚类有效性指标被提出。然而,在某些情况下,存在多个可被选为最终聚类数的选项,而现有大多数研究忽视了这一方面。本研究提出一种基于相关性的模糊聚类有效性指标,即WP(Wiroonsri-Preedasawakul)指标。该指标基于数据点对之间的实际距离与该数据点对调整后质心距离之间的相关性定义。我们评估了该指标的性能,并与包括Xie-Beni、Pakhira-Bandyopadhyay-Maulik、Tang、Wu-Li、广义C及Kwon2在内的多种现有指标进行对比。基于模糊C均值算法,我们在四类数据集(人工数据集、真实数据集、排序模拟数据集及图像数据集)上开展了评估。总体而言,WP指标在准确检测最优聚类数及提供精确次级选项方面优于大多数甚至全部对比指标。此外,即使模糊参数$m$设为较大值,该指标仍保持有效性。本工作使用的R软件包WPfuzzyCVIs可在https://github.com/nwiroonsri/WPfuzzyCVIs获取。