Modeling multiple sampling densities within a hierarchical framework enables borrowing of information across samples. These density random effects can act as kernels in latent variable models to represent exchangeable subgroups or clusters. A key feature of these kernels is the (functional) covariance they induce, which determines how densities are grouped in mixture models. Our motivating problem is clustering chromatin accessibility profiles from high-throughput DNase-seq experiments to detect transcription factor (TF) binding. TF binding typically produces footprint profiles with spatial patterns, creating long-range dependency across genomic locations. Existing nonparametric hierarchical models impose restrictive covariance assumptions and cannot accommodate such dependencies, often leading to biologically uninformative clusters. We propose a nonparametric density kernel flexible enough to capture diverse covariance structures and adaptive to various spatial patterns of TF footprints. The kernel specifies dyadic tree splitting probabilities via a multivariate logit-normal model with a sparse precision matrix. Bayesian inference for latent variable models using this kernel is implemented through Gibbs sampling with Polya-Gamma augmentation. Extensive simulations show that our kernel substantially improves clustering accuracy. We apply the proposed mixture model to DNase-seq data from the ENCODE project, which results in biologically meaningful clusters corresponding to binding events of two common TFs.
翻译:在分层框架中对多个采样密度进行建模,能够实现样本间的信息共享。这些密度随机效应可作为潜变量模型中的核函数,用以表征可交换子组或聚类。此类核函数的关键特征在于其诱导的(函数型)协方差结构,该结构决定了混合模型中密度的分组方式。本文的研究动机源于对高通量DNase-seq实验产生的染色质可及性图谱进行聚类,以检测转录因子结合位点。转录因子结合通常会产生具有空间模式的足迹图谱,从而在基因组位点间形成长程依赖关系。现有非参数分层模型受限于严格的协方差假设,无法适应此类依赖结构,常导致聚类结果缺乏生物学意义。我们提出一种灵活的非参数密度核,既能捕获多样化的协方差结构,又能适应转录因子足迹的各种空间模式。该核通过具有稀疏精度矩阵的多元logit-正态模型来指定二叉分裂概率。采用该核的潜变量模型通过Polya-Gamma增广的吉布斯采样实现贝叶斯推理。大量模拟实验表明,我们的核函数显著提升了聚类精度。我们将所提出的混合模型应用于ENCODE项目的DNase-seq数据,成功获得了对应于两种常见转录因子结合事件的生物学意义聚类。