Missing data in time series is a challenging issue affecting time series analysis. Missing data occurs due to problems like data drops or sensor malfunctioning. Imputation methods are used to fill in these values, with quality of imputation having a significant impact on downstream tasks like classification. In this work, we propose a semi-supervised imputation method, ST-Impute, that uses both unlabeled data along with downstream task's labeled data. ST-Impute is based on sparse self-attention and trains on tasks that mimic the imputation process. Our results indicate that the proposed method outperforms the existing supervised and unsupervised time series imputation methods measured on the imputation quality as well as on the downstream tasks ingesting imputed time series.
翻译:时间序列中的缺失数据是影响时间序列分析的挑战性问题。缺失数据通常由数据丢失或传感器故障等问题引起。插补方法用于填充这些缺失值,且插补质量对分类等下游任务具有重要影响。本文提出了一种半监督插补方法ST-Impute,该方法同时利用未标记数据与下游任务的标记数据。ST-Impute基于稀疏自注意力机制,并通过模拟插补过程的任务进行训练。结果表明,所提方法在插补质量及使用插补时间序列的下游任务上均优于现有的监督式和无监督式时间序列插补方法。