Molecular and genomic technological advancements have greatly enhanced our understanding of biological processes by allowing us to quantify key biological variables such as gene expression, protein levels, and microbiome compositions. These breakthroughs have enabled us to achieve increasingly higher levels of resolution in our measurements, exemplified by our ability to comprehensively profile biological information at the single-cell level. However, the analysis of such data faces several critical challenges: limited number of individuals, non-normality, potential dropouts, outliers, and repeated measurements from the same individual. In this article, we propose a novel method, which we call U-statistic based latent variable (ULV). Our proposed method takes advantage of the robustness of rank-based statistics and exploits the statistical efficiency of parametric methods for small sample sizes. It is a computationally feasible framework that addresses all the issues mentioned above simultaneously. An additional advantage of ULV is its flexibility in modeling various types of single-cell data, including both RNA and protein abundance. The usefulness of our method is demonstrated in two studies: a single-cell proteomics study of acute myelogenous leukemia (AML) and a single-cell RNA study of COVID-19 symptoms. In the AML study, ULV successfully identified differentially expressed proteins that would have been missed by the pseudobulk version of the Wilcoxon rank-sum test. In the COVID-19 study, ULV identified genes associated with covariates such as age and gender, and genes that would be missed without adjusting for covariates. The differentially expressed genes identified by our method are less biased toward genes with high expression levels. Furthermore, ULV identified additional gene pathways likely contributing to the mechanisms of COVID-19 severity.
翻译:分子与基因组学技术的进步极大地增进了我们对生物过程的理解,使我们能够量化基因表达、蛋白质水平和微生物组组成等关键生物学变量。这些突破性进展使得我们能够在测量中获得越来越高的分辨率,例如在单细胞水平上全面分析生物信息的能力。然而,此类数据的分析面临若干关键挑战:个体数量有限、非正态性、潜在的漏失值、异常值以及来自同一个体的重复测量。本文提出了一种新方法,我们称之为基于U统计量的潜变量(ULV)方法。该方法利用了基于秩的统计量的稳健性,并发挥了参数方法在小样本量下的统计效率。它是一个计算上可行的框架,能够同时解决上述所有问题。ULV的另一个优势在于其能够灵活地建模各类单细胞数据,包括RNA和蛋白质丰度。我们通过两项研究证明了该方法的实用性:一项针对急性髓系白血病(AML)的单细胞蛋白质组学研究,以及一项针对COVID-19症状的单细胞RNA研究。在AML研究中,ULV成功识别了伪批量版Wilcoxon秩和检验会遗漏的差异表达蛋白质。在COVID-19研究中,ULV识别了与年龄和性别等协变量相关的基因,以及未调整协变量时会被遗漏的基因。我们的方法所识别的差异表达基因较少偏向于高表达水平的基因。此外,ULV还发现了可能影响COVID-19严重程度机制的额外基因通路。