Clustering is a fundamental learning task widely used as a first step in data analysis. For example, biologists often use cluster assignments to analyze genome sequences, medical records, or images. Since downstream analysis is typically performed at the cluster level, practitioners seek reliable and interpretable clustering models. We propose a new deep-learning framework that predicts interpretable cluster assignments at the instance and cluster levels. First, we present a self-supervised procedure to identify a subset of informative features from each data point. Then, we design a model that predicts cluster assignments and a gate matrix that leads to cluster-level feature selection. We show that the proposed method can reliably predict cluster assignments using synthetic and real data. Furthermore, we verify that our model leads to interpretable results at a sample and cluster level.
翻译:聚类是一种基础的学习任务,广泛用作数据分析的第一步。例如,生物学家常利用聚类结果分析基因组序列、医疗记录或图像。由于下游分析通常在聚类层面进行,研究者需要可靠且可解释的聚类模型。我们提出了一种新的深度学习框架,能够在实例和聚类层面预测可解释的聚类分配。首先,我们设计了一种自监督过程,从每个数据点中识别出信息量丰富的特征子集。随后,我们构建了一个模型,用于预测聚类分配,并生成一个门控矩阵以实现聚类层面的特征选择。实验表明,该方法在合成数据和真实数据上均能可靠地预测聚类分配。此外,我们验证了该模型在样本和聚类层面均能产生可解释的结果。