The fraction of variance explained (FVE) in a linear model quantifies the extent to which predictors account for outcome variability. In high-dimensional settings, where traditional FVE estimators do not apply, modern FVE estimators such as GWASH or linear mix-effect model estimated through the restricted maximum likelihood (LMM-REML) struggle with strong correlation among predictors, often found, for example, in brain imaging data. We propose a decomposition framework that partitions the FVE into two components: a low-dimensional component capturing the strong correlation, estimable by low dimensional methods, and a high-dimensional component with remaining weak correlation, estimable by high dimensional methods. Simulations demonstrate that decomposing dominant principal components (PCs) and estimating the high-dimensional FVE using GWASH or LMM-REML leads to improved bias reduction compared to directly applying standard approaches such as GWASH and LMM-REML. Our method shows consistent performance asymptotically as both the number of predictors and the number of samples increase. We illustrate the method in an analysis of the Adolescent Brain Cognitive Development (ABCD) brain imaging dataset, capturing nuanced heritability signals in the FVE of cognitive measures predicted by high-resolution brain imaging data.
翻译:线性模型中的解释方差比例(FVE)量化了预测变量对结果变异性的解释程度。在高维场景下,传统FVE估计量失效,而现代FVE估计方法(如GWASH或通过限制性最大似然估计的线性混合效应模型(LMM-REML))在预测变量间存在强相关时表现不佳——这一现象常见于脑影像数据。我们提出一种分解框架,将FVE划分为两个分量:捕捉强相关性的低维分量(可通过低维方法估计)与包含剩余弱相关性的高维分量(可通过高维方法估计)。模拟实验表明,相较于直接应用GWASH和LMM-REML等标准方法,对主成分(PCs)进行分解并采用GWASH或LMM-REML估计高维FVE能更有效地降低偏差。随着预测变量数量与样本量同步增加,该方法展现出渐近一致性。我们通过分析青少年大脑认知发展(ABCD)脑影像数据集验证了该方法,从高分辨率脑影像数据预测认知指标的FVE中捕捉到了精细的遗传力信号。