Single-cell sequencing technologies have significantly advanced molecular and cellular biology, offering unprecedented insights into cellular heterogeneity by allowing for the measurement of gene expression at an individual cell level. However, the analysis of such data is challenged by the prevalence of low counts due to dropout events and the skewed nature of the data distribution, which conventional Gaussian factor models struggle to handle effectively. To address these challenges, we propose a novel Bayesian segmented Gaussian copula model to explicitly account for inflation of zero and near-zero counts, and to address the high skewness in the data. By employing a Dirichlet-Laplace prior for each column of the factor loadings matrix, we shrink the loadings of unnecessary factors towards zero, which leads to a simple approach to automatically determine the number of latent factors, and resolve the identifiability issue inherent in factor models due to the rotational invariance of the factor loadings matrix. Through simulation studies, we demonstrate the superior performance of our method over existing approaches in conducting factor analysis on data exhibiting the characteristics of single-cell data, such as excessive low counts and high skewness. Furthermore, we apply the proposed method to a real single-cell RNA-sequencing dataset from a lymphoblastoid cell line, successfully identifying biologically meaningful latent factors and detecting previously uncharacterized cell subtypes.
翻译:单细胞测序技术通过实现单个细胞水平的基因表达测量,在分子与细胞生物学领域取得了显著进展,为细胞异质性研究提供了前所未有的洞察。然而,此类数据分析面临两大挑战:因丢失事件导致的低计数普遍存在,以及数据分布固有的偏斜特性,这使传统高斯因子模型难以有效处理。为应对这些挑战,我们提出一种新型贝叶斯分段高斯连接函数模型,明确考量零计数与近零计数的膨胀现象,并解决数据的高偏度问题。通过为因子载荷矩阵的每一列引入狄利克雷-拉普拉斯先验分布,我们将非必要因子的载荷向零压缩,从而形成一种自动确定潜在因子数量的简便方法,并解决了因子载荷矩阵旋转不变性所导致的因子模型可识别性难题。模拟研究证实,在处理具有单细胞数据特征(如过度低计数与高偏斜度)的数据时,本方法进行因子分析的性能优于现有方法。此外,我们还将所提方法应用于来自淋巴母细胞系的实际单细胞RNA测序数据集,成功识别出具有生物学意义的潜在因子,并检测到先前未被表征的细胞亚型。