Time series clustering is an important data mining task with a wide variety of applications. While most methods focus on time series taking values on the real line, very few works consider functional time series. However, functional objects frequently arise in many fields, such as actuarial science, demography or finance. Functional time series are indexed collections of infinite-dimensional curves viewed as random elements taking values in a Hilbert space. In this paper, the problem of clustering functional time series is addressed. To this aim, a distance between functional time series is introduced and used to construct a clustering procedure. The metric relies on a measure of serial dependence which can be seen as a natural extension of the classical quantile autocorrelation function to the functional setting. Since the dynamics of the series may vary over time, we adopt a fuzzy approach, which enables the procedure to locate each series into several clusters with different membership degrees. The resulting algorithm can group series generated from similar stochastic processes, reaching accurate results with series coming from a broad variety of functional models and requiring minimum hyperparameter tuning. Several simulation experiments show that the method exhibits a high clustering accuracy besides being computationally efficient. Two interesting applications involving high-frequency financial time series and age-specific mortality improvement rates illustrate the potential of the proposed approach.
翻译:时间序列聚类是一项重要的数据挖掘任务,具有广泛的应用场景。尽管大多数方法聚焦于实数值时间序列,但针对函数型时间序列的研究却相对有限。然而,函数型对象在精算科学、人口学或金融学等诸多领域频繁出现。函数型时间序列是无穷维曲线的索引集合,被视为取值于希尔伯特空间的随机元素。本文探讨了函数型时间序列的聚类问题。为此,我们引入了一种函数型时间序列间的距离度量,并基于该度量构建了聚类方法。该度量依赖于序列相关性测度,可视为经典分位数自相关函数在函数型场景中的自然扩展。由于序列的动态特征可能随时间变化,我们采用模糊方法,使得算法能够将每条序列以不同隶属度划分至多个聚类。所提出的算法可对源于相似随机过程的序列进行分组,在覆盖多种函数型模型的序列中取得精确结果,且仅需极少的超参数调优。多项模拟实验表明,该方法在具备高聚类精度的同时兼具计算效率。基于高频金融时间序列和年龄别死亡率改善率的两项实际应用案例,进一步验证了该方法的潜力。