Canonical Correlation Analysis (CCA) is a widespread technique for discovering linear relationships between two sets of variables $X \in \mathbb{R}^{n \times p}$ and $Y \in \mathbb{R}^{n \times q}$. In high dimensions however, standard estimates of the canonical directions cease to be consistent without assuming further structure. In this setting, a possible solution consists in leveraging the presumed sparsity of the solution: only a subset of the covariates span the canonical directions. While the last decade has seen a proliferation of sparse CCA methods, practical challenges regarding the scalability and adaptability of these methods still persist. To circumvent these issues, this paper suggests an alternative strategy that uses reduced rank regression to estimate the canonical directions when one of the datasets is high-dimensional while the other remains low-dimensional. By casting the problem of estimating the canonical direction as a regression problem, our estimator is able to leverage the rich statistics literature on high-dimensional regression and is easily adaptable to accommodate a wider range of structural priors. Our proposed solution maintains computational efficiency and accuracy, even in the presence of very high-dimensional data. We validate the benefits of our approach through a series of simulated experiments and further illustrate its practicality by applying it to three real-world datasets.
翻译:典型相关分析(CCA)是一种广泛应用于发现两组变量$X \in \mathbb{R}^{n \times p}$与$Y \in \mathbb{R}^{n \times q}$间线性关系的技术。然而在高维情形下,若不引入额外结构假设,标准典型方向估计量将不再具有一致性。在此背景下,一种可行的解决方案是利用解的预设稀疏性:仅部分协变量张成典型方向。尽管过去十年见证了稀疏CCA方法的激增,但这些方法在可扩展性与适应性方面仍存在实际挑战。为规避这些问题,本文提出一种替代策略:当其中一个数据集为高维而另一个保持低维时,采用降秩回归来估计典型方向。通过将典型方向估计问题转化为回归问题,我们的估计量能够充分利用高维回归领域丰富的统计学文献,并可轻松适配更广泛的结构先验。即使面对极高维数据,我们提出的方法仍能保持计算效率与精度。我们通过一系列模拟实验验证了该方法的优势,并进一步通过三个真实世界数据集的应用展示了其实用性。