In multivariate time series systems, key insights can be obtained by discovering lead-lag relationships inherent in the data, which refer to the dependence between two time series shifted in time relative to one another, and which can be leveraged for the purposes of control, forecasting or clustering. We develop a clustering-driven methodology for the robust detection of lead-lag relationships in lagged multi-factor models. Within our framework, the envisioned pipeline takes as input a set of time series, and creates an enlarged universe of extracted subsequence time series from each input time series, by using a sliding window approach. We then apply various clustering techniques (e.g, K-means++ and spectral clustering), employing a variety of pairwise similarity measures, including nonlinear ones. Once the clusters have been extracted, lead-lag estimates across clusters are aggregated to enhance the identification of the consistent relationships in the original universe. Since multivariate time series are ubiquitous in a wide range of domains, we demonstrate that our method is not only able to robustly detect lead-lag relationships in financial markets, but can also yield insightful results when applied to an environmental data set.
翻译:在多变量时间序列系统中,通过发现数据中固有的领先滞后关系(即两个时间序列在时间上相互偏移时的依赖关系)可获取关键洞见,这些关系可应用于控制、预测或聚类等任务。我们提出了一种基于聚类的鲁棒性检测方法,专门用于滞后多因子模型中的领先滞后关系识别。在该框架中,所设计的流程以一组时间序列作为输入,通过滑动窗口方法从每个输入时间序列中提取子序列,从而构建一个扩展的时间子序列集合。随后,我们采用多种聚类技术(例如K-means++和谱聚类),并结合包括非线性度量在内的多种成对相似度指标。在完成聚类后,通过跨聚类聚合领先滞后估计值,以增强原始集合中一致性关系的识别能力。鉴于多变量时间序列在众多领域广泛存在,我们证明该方法不仅能够稳健地检测金融市场中的领先滞后关系,在应用于环境数据集时也能产出富有洞察力的结果。