Synthetic Time Series Generation (TSG) is crucial in a range of applications, including data augmentation, anomaly detection, and privacy preservation. Although significant strides have been made in this field, existing methods exhibit three key limitations: (1) They often benchmark against similar model types, constraining a holistic view of performance capabilities. (2) The use of specialized synthetic and private datasets introduces biases and hampers generalizability. (3) Ambiguous evaluation measures, often tied to custom networks or downstream tasks, hinder consistent and fair comparison. To overcome these limitations, we introduce \textsf{TSGBench}, the inaugural TSG Benchmark, designed for a unified and comprehensive assessment of TSG methods. It comprises three modules: (1) a curated collection of publicly available, real-world datasets tailored for TSG, together with a standardized preprocessing pipeline; (2) a comprehensive evaluation measures suite including vanilla measures, new distance-based assessments, and visualization tools; (3) a pioneering generalization test rooted in Domain Adaptation (DA), compatible with all methods. We have conducted extensive experiments across ten real-world datasets from diverse domains, utilizing ten advanced TSG methods and twelve evaluation measures, all gauged through \textsf{TSGBench}. The results highlight its remarkable efficacy and consistency. More importantly, \textsf{TSGBench} delivers a statistical breakdown of method rankings, illuminating performance variations across different datasets and measures, and offering nuanced insights into the effectiveness of each method.
翻译:合成时间序列生成(TSG)在数据增强、异常检测和隐私保护等一系列应用中至关重要。尽管该领域已取得显著进展,但现有方法存在三个关键局限性:(1)它们通常针对相似模型类型进行基准测试,限制了性能能力的整体视角;(2)使用专门的合成和私有数据集引入了偏差并损害了泛化能力;(3)模糊的评估指标(通常与自定义网络或下游任务相关联)阻碍了一致且公平的比较。为克服这些局限性,我们提出 \textsf{TSGBench}——首个 TSG 基准,旨在对 TSG 方法进行统一且全面的评估。它包含三个模块:(1)精心整理的、适用于 TSG 的公开真实世界数据集集合,以及标准化预处理流程;(2)一套全面的评估指标,包括原始指标、新的基于距离的评估和可视化工具;(3)一种基于领域自适应(DA)的开创性泛化测试,兼容所有方法。我们利用 \textsf{TSGBench},在来自不同领域的十个真实世界数据集上,使用十种先进 TSG 方法和十二种评估指标进行了广泛实验。结果凸显了其显著的有效性和一致性。更为重要的是,\textsf{TSGBench} 提供了方法排名的统计分解,揭示了不同数据集和指标下的性能差异,并提供了关于每种方法有效性的细致洞察。