Self-supervised learning approaches provide a promising direction for clustering multivariate time-series data. However, real-world time-series data often include missing values, and the existing approaches require imputing missing values before clustering, which may cause extensive computations and noise and result in invalid interpretations. To address these challenges, we present a Self-supervised Learning-based Approach to Clustering multivariate Time-series data with missing values (SLAC-Time). SLAC-Time is a Transformer-based clustering method that uses time-series forecasting as a proxy task for leveraging unlabeled data and learning more robust time-series representations. This method jointly learns the neural network parameters and the cluster assignments of the learned representations. It iteratively clusters the learned representations with the K-means method and then utilizes the subsequent cluster assignments as pseudo-labels to update the model parameters. To evaluate our proposed approach, we applied it to clustering and phenotyping Traumatic Brain Injury (TBI) patients in the TRACK-TBI dataset. Our experiments demonstrate that SLAC-Time outperforms the baseline K-means clustering algorithm in terms of silhouette coefficient, Calinski Harabasz index, Dunn index, and Davies Bouldin index. We identified three TBI phenotypes that are distinct from one another in terms of clinically significant variables as well as clinical outcomes, including the Extended Glasgow Outcome Scale (GOSE) score, Intensive Care Unit (ICU) length of stay, and mortality rate. The experiments show that the TBI phenotypes identified by SLAC-Time can be potentially used for developing targeted clinical trials and therapeutic strategies.
翻译:自监督学习方法为多变量时间序列数据的聚类提供了有前景的方向。然而,真实世界的时间序列数据通常包含缺失值,现有方法需要在聚类前对缺失值进行插补,这可能导致大量计算和噪声,并产生无效的解释。为解决这些挑战,我们提出了一种基于自监督学习的多变量时间序列数据含缺失值聚类方法(SLAC-Time)。SLAC-Time是一种基于Transformer的聚类方法,利用时间序列预测作为代理任务,以充分利用无标注数据并学习更鲁棒的时间序列表示。该方法联合学习神经网络参数与所学表示的聚类分配,通过K-means方法迭代地对所学表示进行聚类,并利用后续的聚类分配作为伪标签更新模型参数。为评估所提方法,我们将其应用于TRACK-TBI数据集中创伤性脑损伤(TBI)患者的聚类与表型分析。实验结果表明,SLAC-Time在轮廓系数、Calinski-Harabasz指数、Dunn指数和Davies-Bouldin指数上均优于基线K-means聚类算法。我们识别出三种TBI表型,这些表型在临床重要变量以及临床结局(包括扩展格拉斯哥结局量表(GOSE)评分、重症监护室(ICU)住院时间和死亡率)上彼此显著不同。实验表明,SLAC-Time识别的TBI表型可能用于开发靶向临床试验和治疗策略。