Canonical correlation analysis (CCA) is a technique for finding correlated sets of features between two datasets. In this paper, we propose a novel extension of CCA to the online, streaming data setting: Sliding Window Informative Canonical Correlation Analysis (SWICCA). Our method uses a streaming principal component analysis (PCA) algorithm as a backend and uses these outputs combined with a small sliding window of samples to estimate the CCA components in real time. We motivate and describe our algorithm, provide numerical simulations to characterize its performance, and provide a theoretical performance guarantee. The SWICCA method is applicable and scalable to extremely high dimensions, and we provide a real-data example that demonstrates this capability.
翻译:典型相关分析(CCA)是一种用于发现两个数据集之间特征相关关系的技术。本文提出一种面向在线流数据场景的CCA新型扩展方法:滑动窗口信息典型相关分析(SWICCA)。该方法以流式主成分分析(PCA)算法为后端,结合小型滑动窗口样本的实时输出,实现CCA分量的在线估计。我们对算法进行了理论阐释与描述,通过数值仿真刻画其性能表现,并提供了理论性能保证。SWICCA方法适用于超高维场景且具备可扩展性,文中通过真实数据示例验证了其实际应用能力。