Text document clustering can play a vital role in organizing and handling the everincreasing number of text documents. Uninformative and redundant features included in large text documents reduce the effectiveness of the clustering algorithm. Feature selection (FS) is a well-known technique for removing these features. Since FS can be formulated as an optimization problem, various meta-heuristic algorithms have been employed to solve it. Teaching-Learning-Based Optimization (TLBO) is a novel meta-heuristic algorithm that benefits from the low number of parameters and fast convergence. A hybrid method can simultaneously benefit from the advantages of TLBO and tackle the possible entrapment in the local optimum. By proposing a hybrid of TLBO, Grey Wolf Optimizer (GWO), and Genetic Algorithm (GA) operators, this paper suggests a filter-based FS algorithm (TLBO-GWO). Six benchmark datasets are selected, and TLBO-GWO is compared with three recently proposed FS algorithms with similar approaches, the main TLBO and GWO. The comparison is conducted based on clustering evaluation measures, convergence behavior, and dimension reduction, and is validated using statistical tests. The results reveal that TLBO-GWO can significantly enhance the effectiveness of the text clustering technique (K-means).
翻译:文本聚类在组织和处理日益增长的文本文档中发挥着关键作用。大规模文本中包含的无信息特征和冗余特征会降低聚类算法的有效性。特征选择(FS)是去除这些特征的经典技术。由于FS可被建模为优化问题,多种元启发式算法已被用于求解该问题。教学优化算法(TLBO)作为一种新型元启发式算法,具有参数少、收敛快等优势。混合方法能同时利用TLBO的优势并避免其陷入局部最优。本文通过融合TLBO、灰狼优化(GWO)及遗传算法(GA)算子,提出一种基于过滤的FS算法(TLBO-GWO)。选取六个基准数据集,将TLBO-GWO与三种采用相似思路的最新FS算法、原始TLBO及GWO进行对比。基于聚类评估指标、收敛行为和降维效果进行对比分析,并通过统计检验验证结果。实验表明,TLBO-GWO能显著提升文本聚类技术(K-means)的有效性。