Since their introduction in the last few years, conditional generative models have seen remarkable achievements. However, they often need the use of large amounts of labelled information. By using unsupervised conditional generation in conjunction with a clustering inference network, ClusterGAN has recently been able to achieve impressive clustering results. Since the real conditional distribution of data is ignored, the clustering inference network can only achieve inferior clustering performance by considering only uniform prior based generative samples. However, the true distribution is not necessarily balanced. Consequently, ClusterGAN fails to produce all modes, which results in sub-optimal clustering inference network performance. So, it is important to learn the prior, which tries to match the real distribution in an unsupervised way. In this paper, we propose self-augmentation information maximization improved ClusterGAN (SIMI-ClusterGAN) to learn the distinctive priors from the data directly. The proposed SIMI-ClusterGAN consists of four deep neural networks: self-augmentation prior network, generator, discriminator and clustering inference network. The proposed method has been validated using seven benchmark data sets and has shown improved performance over state-of-the art methods. To demonstrate the superiority of SIMI-ClusterGAN performance on imbalanced dataset, we have discussed two imbalanced conditions on MNIST datasets with one-class imbalance and three classes imbalanced cases. The results highlight the advantages of SIMI-ClusterGAN.
翻译:自条件生成模型引入以来,近年来取得了显著成就。然而,它们通常需要使用大量标注信息。通过将无监督条件生成与聚类推理网络相结合,ClusterGAN最近在聚类任务上取得了令人瞩目的成果。由于忽略了数据的真实条件分布,聚类推理网络仅考虑基于均匀先验的生成样本,导致其聚类性能较差。然而,真实分布未必是均衡的。因此,ClusterGAN无法生成所有模态,使得聚类推理网络性能次优。所以,以无监督方式学习与真实分布匹配的先验至关重要。本文提出自增强信息最大化改进ClusterGAN(SIMI-ClusterGAN),直接从数据中学习独特的先验分布。所提出的SIMI-ClusterGAN包含四个深度神经网络:自增强先验网络、生成器、判别器和聚类推理网络。该方法在七个基准数据集上进行了验证,其性能相较于现有最优方法有所提升。为展示SIMI-ClusterGAN在不平衡数据集上的优越性,我们在MNIST数据集中讨论了单类不平衡和三类别不平衡两种不平衡情形。实验结果凸显了SIMI-ClusterGAN的优势。