This paper presents a comparative analysis of different optimization techniques for the K-means algorithm in the context of big data. K-means is a widely used clustering algorithm, but it can suffer from scalability issues when dealing with large datasets. The paper explores different approaches to overcome these issues, including parallelization, approximation, and sampling methods. The authors evaluate the performance of these techniques on various benchmark datasets and compare them in terms of speed, quality of clustering, and scalability according to the LIMA dominance criterion. The results show that different techniques are more suitable for different types of datasets and provide insights into the trade-offs between speed and accuracy in K-means clustering for big data. Overall, the paper offers a comprehensive guide for practitioners and researchers on how to optimize K-means for big data applications.
翻译:本文针对大数据背景下K-means算法的不同优化技术进行了比较分析。K-means是一种广泛使用的聚类算法,但在处理大规模数据集时存在可扩展性问题。本文探讨了克服这些问题的不同方法,包括并行化、近似和采样技术。作者在多个基准数据集上评估了这些技术的性能,并根据LIMA支配准则在速度、聚类质量和可扩展性方面进行了比较。结果表明,不同技术适用于不同类型的数据集,并揭示了在大数据K-means聚类中速度与精度之间的权衡关系。总体而言,本文为实践者和研究人员优化大数据应用中的K-means算法提供了全面指南。