The importance of variable selection for clustering has been recognized for some time, and mixture models are well-established as a statistical approach to clustering. Yet, the literature on variable selection in model-based clustering remains largely rooted in the assumption of Gaussian clusters. Unsurprisingly, variable selection algorithms based on this assumption tend to break down in the presence of cluster skewness. A novel variable selection algorithm is presented that utilizes the Manly transformation mixture model to select variables based on their ability to separate clusters, and is effective even when clusters depart from the Gaussian assumption. The proposed approach, which is implemented within the R package vscc, is compared to existing variable selection methods -- including an existing method that can account for cluster skewness -- using simulated and real datasets
翻译:变量选择对于聚类的重要性早已得到公认,而混合模型作为聚类的统计方法也已相当成熟。然而,基于模型的聚类变量选择研究仍主要植根于高斯聚类的假设。基于该假设的变量选择算法在存在聚类偏态时往往会失效,这并不令人意外。本文提出了一种新型变量选择算法,该算法利用曼利变换混合模型,根据变量分离聚类的能力进行选择,即便在聚类偏离高斯假设的情况下也能有效工作。所提出的方法已在R包vscc中实现,并通过模拟数据和真实数据集与现有变量选择方法(包括一种能够处理聚类偏态的现有方法)进行了比较。