Model-based clustering of moderate or large dimensional data is notoriously difficult. We propose a model for simultaneous dimensionality reduction and clustering by assuming a mixture model for a set of latent scores, which are then linked to the observations via a Gaussian latent factor model. This approach was recently investigated by Chandra et al. (2020). The authors use a factor-analytic representation and assume a mixture model for the latent factors. However, performance can deteriorate in the presence of model misspecification. Assuming a repulsive point process prior for the component-specific means of the mixture for the latent scores is shown to yield a more robust model that outperforms the standard mixture model for the latent factors in several simulated scenarios. To favor well-separated clusters of data, the repulsive point process must be anisotropic, and its density should be tractable for efficient posterior inference. We address these issues by proposing a general construction for anisotropic determinantal point processes.
翻译:针对中等或高维数据的基于模型的聚类通常非常困难。我们通过假设一组潜在得分的混合模型,并借助高斯潜在因子模型将这些得分与观测值相关联,提出了一个同时实现降维和聚类的模型。该方法最近由Chandra等人(2020)研究过。作者采用因子分析表示,并假设潜在因子服从混合模型。然而,在模型误设的情况下,性能可能会下降。研究表明,对潜在得分混合的组件特定均值施加排斥点过程先验,能够产生更稳健的模型,并在若干模拟场景中优于标准潜在因子混合模型。为了促进数据形成良好分离的聚类,排斥点过程必须是各向异性的,且其密度应易于处理以实现高效的后验推断。我们通过提出各向异性行列式点过程的通用构造来解决这些问题。