Spatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the middle temporal module catches inter-frame correlations. While the mainstream methods employ recurrent units to capture long-term temporal dependencies, they suffer from low computational efficiency due to their unparallelizable architectures. To parallelize the temporal module, we propose the Temporal Attention Unit (TAU), which decomposes the temporal attention into intra-frame statical attention and inter-frame dynamical attention. Moreover, while the mean squared error loss focuses on intra-frame errors, we introduce a novel differential divergence regularization to take inter-frame variations into account. Extensive experiments demonstrate that the proposed method enables the derived model to achieve competitive performance on various spatiotemporal prediction benchmarks.
翻译:时空预测学习旨在通过从历史帧中学习来生成未来帧。本文对现有方法进行探究,并提出了一个通用的时空预测学习框架,其中空间编码器和解码器负责捕捉帧内特征,而中间的时间模块则捕捉帧间相关性。尽管主流方法采用循环单元来捕获长期时间依赖性,但其不可并行化的架构导致计算效率低下。为了实现时间模块的并行化,我们提出了时间注意力单元(Temporal Attention Unit, TAU),该单元将时间注意力分解为帧内静态注意力和帧间动态注意力。此外,针对均方误差损失仅聚焦于帧内误差的问题,我们引入了一种新颖的差分散度正则化方法,以考虑帧间变化。大量实验表明,所提方法使得衍生模型在多种时空预测基准上取得了具有竞争力的性能。