Spatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the middle temporal module catches inter-frame correlations. While the mainstream methods employ recurrent units to capture long-term temporal dependencies, they suffer from low computational efficiency due to their unparallelizable architectures. To parallelize the temporal module, we propose the Temporal Attention Unit (TAU), which decomposes the temporal attention into intra-frame statical attention and inter-frame dynamical attention. Moreover, while the mean squared error loss focuses on intra-frame errors, we introduce a novel differential divergence regularization to take inter-frame variations into account. Extensive experiments demonstrate that the proposed method enables the derived model to achieve competitive performance on various spatiotemporal prediction benchmarks.
翻译:时空预测学习旨在通过学习历史帧来生成未来帧。本文在考察现有方法的基础上,提出了一个通用的时空预测学习框架,其中空间编码器和解码器负责捕捉帧内特征,中间的时间模块负责捕捉帧间相关性。尽管主流方法采用循环单元来捕获长期时间依赖性,但由于其不可并行化的架构,导致计算效率低下。为并行化时间模块,我们提出了时空注意力单元(TAU),该单元将时间注意力分解为帧内静态注意力和帧间动态注意力。此外,针对均方误差损失侧重于帧内误差的问题,我们引入了一种新颖的差分散度正则化方法,以考虑帧间变化。大量实验表明,所提出的方法使模型在多种时空预测基准上取得了具有竞争力的性能。