Recent advances in engineering technologies have enabled the collection of a large number of longitudinal features. This wealth of information presents unique opportunities for researchers to investigate the complex nature of diseases and uncover underlying disease mechanisms. However, analyzing such kind of data can be difficult due to its high dimensionality, heterogeneity and computational challenges. In this paper, we propose a Bayesian nonparametric mixture model for clustering high-dimensional mixed-type (e.g., continuous, discrete and categorical) longitudinal features. We employ a sparse factor model on the joint distribution of random effects and the key idea is to induce clustering at the latent factor level instead of the original data to escape the curse of dimensionality. The number of clusters is estimated through a Dirichlet process prior. An efficient Gibbs sampler is developed to estimate the posterior distribution of the model parameters. Analysis of real and simulated data is presented and discussed. Our study demonstrates that the proposed model serves as a useful analytical tool for clustering high-dimensional longitudinal data.
翻译:近年来工程技术的进步使得大规模纵向特征的收集成为可能。这些丰富的数据为研究人员探究疾病的复杂性质并揭示潜在疾病机制提供了独特机遇。然而,由于此类数据的高维度、异质性及计算挑战,其分析过程可能面临困难。本文提出一种贝叶斯非参数混合模型,用于对高维混合类型(如连续、离散和分类)的纵向特征进行聚类。我们在随机效应的联合分布中采用稀疏因子模型,其核心思想是通过在潜在因子层面而非原始数据层面诱导聚类,以规避维数灾难。簇的数量通过狄利克雷过程先验进行估计。我们开发了一种高效的吉布斯采样器来估计模型参数的后验分布,并对真实数据与模拟数据的分析结果进行了呈现与讨论。研究表明,所提出的模型可作为高维纵向数据聚类的有效分析工具。