Recently, remarkable progress has been made over large language models (LLMs), demonstrating their unprecedented capability in varieties of natural language tasks. However, completely training a large general-purpose model from the scratch is challenging for time series analysis, due to the large volumes and varieties of time series data, as well as the non-stationarity that leads to concept drift impeding continuous model adaptation and re-training. Recent advances have shown that pre-trained LLMs can be exploited to capture complex dependencies in time series data and facilitate various applications. In this survey, we provide a systematic overview of existing methods that leverage LLMs for time series analysis. Specifically, we first state the challenges and motivations of applying language models in the context of time series as well as brief preliminaries of LLMs. Next, we summarize the general pipeline for LLM-based time series analysis, categorize existing methods into different groups (i.e., direct query, tokenization, prompt design, fine-tune, and model integration), and highlight the key ideas within each group. We also discuss the applications of LLMs for both general and spatial-temporal time series data, tailored to specific domains. Finally, we thoroughly discuss future research opportunities to empower time series analysis with LLMs.
翻译:近年来,大型语言模型取得了显著进展,展示了其在各种自然语言任务中前所未有的能力。然而,由于时间序列数据量庞大、类型多样,且非平稳性导致概念漂移阻碍模型的持续适应与重新训练,从头开始完整训练一个通用大型模型对时间序列分析而言颇具挑战。近期研究表明,预训练的大型语言模型可用于捕捉时间序列数据中的复杂依赖关系,并促进各类应用发展。本综述系统梳理了现有利用大型语言模型进行时间序列分析的方法。具体而言,我们首先阐述了将语言模型应用于时间序列场景的挑战与动机,并简要介绍了大型语言模型的基础知识。接着,我们总结了基于大型语言模型的时间序列分析通用流程,将现有方法归为不同类别(即直接查询、分词化、提示设计、微调和模型集成),并强调了每类方法的核心思想。我们还讨论了大型语言模型在通用时间序列数据及面向特定领域的时空时间序列数据中的应用。最后,我们深入探讨了利用大型语言模型赋能时间序列分析的未来研究机遇。