As time-series applications grow larger, there is increasing demand for symbolic representations that are compact, accurate, and scalable across many signals and computing resources. Current ABBA-based symbolic approximation methods produce high-quality, shape-preserving representations, but they handle each time series separately and sequentially. This means they do not ensure consistent symbols across different series and cannot fully exploit modern multicore systems and distributed-memory systems. This paper presents a joint symbolic time-series approximation method for large-scale time series. The proposed method decouples local compression from global digitization: (i) time series are partitioned into independent domains that can be compressed in parallel, and (ii) the resulting pieces are digitized using a shared global dictionary. To further improve scalability, we introduce a two-stage parallel digitization scheme, in which aggregation is first performed locally and then merged globally without requiring a full-data reassignment step. Extensive experiments on time-series datasets and large synthetic benchmarks show that our approach maintains competitive reconstruction quality while substantially reducing runtime. These results show that joint symbolic approximation can serve as an efficient, high-level parallel tool for analyzing large-scale temporal data.
翻译:随着时间序列应用规模的日益增长,对能够跨多个信号和计算资源实现紧凑、精确且可扩展的符号表示的需求不断增加。当前基于ABBA的符号近似方法能生成高质量且保留形状特征的表示,但需对每个时间序列分别进行顺序处理。这意味着无法保证不同序列间符号的一致性,也无法充分利用现代多核系统与分布式存储系统的计算能力。本文提出一种面向大规模时间序列的联合符号时序近似方法。该方法将局部压缩与全局数字化解耦:(i) 将时间序列划分为可在并行环境中独立压缩的域单元;(ii) 基于共享全局词典对压缩后的片段进行数字化。为进一步提升可扩展性,我们引入两阶段并行数字化方案:先执行局部聚合,再通过全局合并完成最终数字化,无需完整数据重分配步骤。在时间序列数据集和大型合成基准上的广泛实验表明,本方法在保持竞争性重构质量的同时显著缩短了运行时间。这些结果表明,联合符号近似可作为分析大规模时序数据的高效高层并行工具。