Despite tremendous advancements in Artificial Intelligence, learning from large sets of data in an unsupervised manner remains a significant challenge. Classical clustering algorithms often fail to discover complex dependencies in large datasets, especially considering sparse, high-dimensional spaces. However, deep learning techniques proved to be successful when dealing with large quantities of data, efficiently reducing their dimensionality without losing track of underlying information. Several interesting advancements have already been made to combine deep learning and clustering. Still, the idea of enhancing the clustering results by combining multiple views of the data generated by deep neural networks appears to be insufficiently explored yet. This paper aims to investigate this direction and bridge the gap between deep neural networks, clustering techniques and ensemble learning methods. To achieve this goal, we propose a novel deep clustering ensemble method - Snapshot Spectral Clustering, designed to maximize the gain from combining multiple data views while minimizing the computational costs of creating the ensemble. Comparative analysis and experiments described in this paper prove the proposed concept, while the conducted hyperparameter study provides a valuable intuition to follow when selecting proper values.
翻译:尽管人工智能取得了巨大进步,但从大规模数据中以无监督方式学习仍是一项重大挑战。经典聚类算法在处理大型数据集中的复杂依赖关系时常常失败,尤其是在稀疏的高维空间中。然而,深度学习技术在处理大量数据时已被证明是成功的,它能在保持底层信息的同时有效降低数据维度。尽管已有一些将深度学习与聚类相结合的有趣进展,但通过利用深度神经网络生成的多个数据视图来增强聚类结果的想法似乎仍未得到充分探索。本文旨在研究这一方向,并弥合深度神经网络、聚类技术和集成学习方法之间的差距。为实现这一目标,我们提出了一种新颖的深度聚类集成方法——快照谱聚类,旨在最大化从组合多个数据视图中获得的收益,同时最小化创建集成的计算成本。本文中的比较分析和实验证明了所提出的概念,而超参数研究则为选择适当参数提供了宝贵的直观指导。