Regionalization aims to partition a spatial domain into contiguous regions that share similar characteristics, enabling more effective spatial analysis, policy making, and resource management. Existing approaches for spatial regionalization typically rely on static spatial snapshots rather than evolving time series. Meanwhile, most time series clustering methods ignore spatial structure or enforce spatial continuity through ad hoc regularization, constraining the number of inferred regions a priori either explicitly or implicitly. Utilizing the minimum description length principle from information theory, here we propose an efficient and fully nonparametric framework for the regionalization of spatial time series. Our method jointly infers a spatial partition along with a set of representative time series archetypes ("drivers") that best compress a spatiotemporal dataset, with a runtime log-linear in the number of time series. We demonstrate that this method can accurately recover planted regional structure and drivers in synthetic time series, and can extract meaningful structural regularities in large-scale empirical air quality and vegetation index records. Our method provides a principled and scalable framework for spatially contiguous partitioning, allowing interpretable temporal patterns and homogeneous regions to emerge directly from the data itself.
翻译:区域化旨在将空间域划分为具有相似特征的连续区域,从而提升空间分析、政策制定及资源管理的有效性。现有空间区域化方法通常依赖静态空间快照而非动态演化的时间序列。与此同时,多数时间序列聚类方法忽略空间结构,或通过临时正则化强制空间连续性,显式或隐式地预先约束推断区域的数量。基于信息论中的最小描述长度原理,本文提出一种高效且完全非参数化的空间时间序列区域化框架。该方法联合推断空间划分与一组代表性时间序列原型("驱动因子"),以最优压缩时空数据集,其运行时间与时间序列数量呈对数线性关系。我们证明该方法能精确恢复合成时间序列中的植入区域结构与驱动因子,并在大规模空气质量与植被指数实测记录中提取有意义的结构规律。该方法为空间连续划分提供了原理性且可扩展的框架,使可解释的时间模式与同质区域能够直接从数据中自然涌现。