In this work, we present a novel method for divisive hierarchical variable clustering. A cluster is a group of elements that exhibit higher similarity among themselves than to elements outside this cluster. The correlation coefficient serves as a natural measure to assess the similarity of variables. This means that in a correlation matrix, a cluster is represented by a block of variables with greater internal than external correlation. Our approach provides a nonparametric solution to identify such block structures in the correlation matrix using singular vectors of the underlying data matrix. When divisively clustering $p$ variables, there are $2^{p-1}$ possible splits. Using the singular vectors for cluster identification, we can effectively reduce these number to at most $p(p-1)$, thereby making it computationally efficient. We elaborate on the methodology and outline the incorporation of dissimilarity measures and linkage functions to assess distances between clusters. Additionally, we demonstrate that these distances are ultrametric, ensuring that the resulting hierarchical cluster structure can be uniquely represented by a dendrogram, with the heights of the dendrogram being interpretable. To validate the efficiency of our method, we perform simulation studies and analyze real world data on personality traits and cognitive abilities. Supplementary materials for this article can be accessed online.
翻译:本文提出了一种新颖的变量分裂式层次聚类方法。聚类是指一组元素,其内部相似度高于与外部元素之间的相似度。相关系数是评估变量相似度的自然度量。这意味着在相关矩阵中,聚类表现为一个变量块,其内部相关性大于外部相关性。我们的方法提供了一种非参数解决方案,利用底层数据矩阵的奇异向量来识别相关矩阵中的这种块结构。当对p个变量进行分裂式聚类时,存在2^(p-1)种可能的划分方式。通过使用奇异向量进行聚类识别,我们可以将这一数量有效减少至最多p(p-1)个,从而提高了计算效率。我们详细阐述了该方法的原理,并说明了如何整合相异度度量和链接函数来评估聚类间的距离。此外,我们证明了这些距离具有超度量性质,确保由此产生的层次聚类结构可以唯一地通过树状图表示,且树状图的高度具有可解释性。为验证该方法的有效性,我们进行了模拟研究,并分析了关于人格特质和认知能力的真实世界数据。本文的补充材料可在网上获取。