While time series classification and forecasting problems have been extensively studied, the cases of noisy time series data with arbitrary time sequence lengths have remained challenging. Each time series instance can be thought of as a sample realization of a noisy dynamical model, which is characterized by a continuous stochastic process. For many applications, the data are mixed and consist of several types of noisy time series sequences modeled by multiple stochastic processes, making the forecasting and classification tasks even more challenging. Instead of regressing data naively and individually to each time series type, we take a latent variable model approach using a mixtured Gaussian processes with learned spectral kernels. More specifically, we auto-assign each type of noisy time series data a signature vector called its motion code. Then, conditioned on each assigned motion code, we infer a sparse approximation of the corresponding time series using the concept of the most informative timestamps. Our unmixing classification approach involves maximizing the likelihood across all the mixed noisy time series sequences of varying lengths. This stochastic approach allows us to learn not only within a single type of noisy time series data but also across many underlying stochastic processes, giving us a way to learn multiple dynamical models in an integrated and robust manner. The different learned latent stochastic models allow us to generate specific sub-type forecasting. We provide several quantitative comparisons demonstrating the performance of our approach.
翻译:尽管时间序列分类与预测问题已得到广泛研究,但具有任意时间序列长度的含噪数据仍具有挑战性。每个时间序列实例可视为一个含噪动态模型的样本实现,该模型由连续随机过程刻画。在许多应用中,数据由多种随机过程建模的含噪时间序列混合而成,这使得预测和分类任务更为复杂。我们未采用对每种时间序列类型进行简单独立回归的方法,而是基于隐变量模型框架,利用含学习谱核的混合高斯过程。具体而言,我们为每类含噪时间序列数据自动分配一个称为运动编码的签名向量。随后,基于每个分配的运动编码,我们利用最具信息量的时间戳概念,推断对应时间序列的稀疏近似。这种解混分类方法通过最大化所有长度不一的混合含噪时间序列的似然函数实现。该随机方法不仅允许我们学习单一类型的含噪时间序列数据,还能跨多个底层随机过程进行学习,从而以集成鲁棒的方式学习多种动态模型。学习到的不同潜在随机模型使我们能够生成特定子类型的预测。我们提供了多项定量比较结果,验证了该方法的性能。