Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with clustered data. There are three commonly used Pearson correlation coefficients (total, between-, and within-cluster), which together provide an enriched perspective of the correlation. However, these Pearson correlation coefficients are sensitive to extreme values and skewed distributions. They also depend on the scale of the data and are not applicable to ordered categorical data. Current non-parametric measures for clustered data are only for the total correlation. Here we define population parameters for the between- and within-cluster Spearman rank correlations. The definitions are natural extensions of the Pearson between- and within-cluster correlations to the rank scale. We show that the total Spearman rank correlation approximates a weighted sum of the between- and within-cluster Spearman rank correlations, where the weights are functions of rank intraclass correlations of the two random variables. We also discuss the equivalence between the within-cluster Spearman rank correlation and the covariate-adjusted partial Spearman rank correlation. Furthermore, we describe estimation and inference for the three Spearman rank correlations, conduct simulations to evaluate the performance of our estimators, and illustrate their use with data from a longitudinal biomarker study and a clustered randomized trial.
翻译:聚类数据在实践中很常见。当受试者被重复测量,或受试者嵌套于群组(如家庭、学校)中时,便产生了聚类现象。评估聚类数据中两个变量之间的相关性通常具有重要意义。有三种常用的皮尔逊相关系数(总体、集群间和集群内),它们共同提供了相关性的丰富视角。然而,这些皮尔逊相关系数对极端值和偏态分布敏感,且依赖于数据的尺度,不适用于有序分类数据。当前针对聚类数据的非参数度量仅适用于总体相关性。本文定义了集群间和集群内Spearman秩相关的总体参数。这些定义是将皮尔逊集群间和集群内相关性自然推广至秩尺度。我们证明,总体Spearman秩相关近似等于集群间和集群内Spearman秩相关的加权和,其中权重是两个随机变量的秩内相关系数的函数。我们还讨论了集群内Spearman秩相关与协变量调整后的偏Spearman秩相关之间的等价性。此外,我们描述了这三种Spearman秩相关的估计与推断方法,通过模拟评估了估计量的性能,并利用一项纵向生物标志物研究和一项集群随机试验的数据说明了其应用。