Biclustering is a powerful unsupervised learning technique for simultaneously identifying coherent subsets of rows and columns in a data matrix, thus revealing local patterns that may not be apparent in global analyses. However, most biclustering methods are developed for continuous data and are not applicable for binary datasets such as single-nucleotide polymorphism (SNP) or protein-protein interaction (PPI) data. Existing biclustering algorithms for binary data often struggle to recover biclustering patterns under noise, face scalability issues, and/or bias the final results towards biclusters of a particular size or characteristic. We propose a Bayesian method for biclustering binary datasets called Binary Spike-and-Slab Lasso Biclustering (BiSSLB). Our method is robust to noise and allows for overlapping biclusters of various sizes without prior knowledge of the noise level or bicluster characteristics. BiSSLB is based on a logistic matrix factorization model with spike-and-slab priors on the latent spaces. We further incorporate an Indian Buffet Process (IBP) prior to automatically determine the number of biclusters from the data. We develop a novel coordinate ascent algorithm with proximal steps which allows for scalable computation. The performance of our proposed approach is assessed through simulations and two real applications on HapMap SNP and Homo Sapiens PPI data, where BiSSLB is shown to outperform other state-of-the-art binary biclustering methods when the data is very noisy.
翻译:摘要:双聚类是一种强大的无监督学习技术,可同时识别数据矩阵中行和列的一致性子集,从而揭示在全局分析中可能不明显的局部模式。然而,大多数双聚类方法是为连续数据设计的,不适用于二元数据集,例如单核苷酸多态性(SNP)或蛋白质-蛋白质相互作用(PPI)数据。现有的二元数据双聚类算法在噪声下难以恢复双聚类模式,面临可扩展性问题,或使最终结果偏向于特定大小或特征的双聚类。我们提出了一种名为二元尖峰-板状套索双聚类(BiSSLB)的贝叶斯方法,用于对二元数据集进行双聚类。该方法对噪声具有鲁棒性,允许不同大小的双聚类重叠,无需预先了解噪声水平或双聚类特征。BiSSLB基于逻辑矩阵分解模型,并在潜在空间上采用尖峰-板状先验。我们进一步引入印度自助餐过程(IBP)先验,以自动从数据中确定双聚类的数量。我们开发了一种新颖的坐标上升算法,结合近端步骤,实现了可扩展的计算。通过模拟实验以及对HapMap SNP和智人PPI数据的两个实际应用,评估了所提方法的性能,结果表明在数据高度噪声的情况下,BiSSLB优于其他最先进的二元双聚类方法。