We address the video prediction task by putting forth a novel model that combines (i) our recently proposed hierarchical residual vector quantized variational autoencoder (HR-VQVAE), and (ii) a novel spatiotemporal PixelCNN (ST-PixelCNN). We refer to this approach as a sequential hierarchical residual learning vector quantized variational autoencoder (S-HR-VQVAE). By leveraging the intrinsic capabilities of HR-VQVAE at modeling still images with a parsimonious representation, combined with the ST-PixelCNN's ability at handling spatiotemporal information, S-HR-VQVAE can better deal with chief challenges in video prediction. These include learning spatiotemporal information, handling high dimensional data, combating blurry prediction, and implicit modeling of physical characteristics. Extensive experimental results on the KTH Human Action and Moving-MNIST tasks demonstrate that our model compares favorably against top video prediction techniques both in quantitative and qualitative evaluations despite a much smaller model size. Finally, we boost S-HR-VQVAE by proposing a novel training method to jointly estimate the HR-VQVAE and ST-PixelCNN parameters.
翻译:我们提出了一种结合(i)近期提出的分层残差向量量化变分自编码器(HR-VQVAE)与(ii)新型时空像素卷积神经网络(ST-PixelCNN)的模型,以解决视频预测任务。该方法被命名为序列化分层残差学习向量量化变分自编码器(S-HR-VQVAE)。通过利用HR-VQVAE在静态图像建模中凭借精简表征的内在能力,并结合ST-PixelCNN处理时空信息的优势,S-HR-VQVAE能更有效地应对视频预测中的核心挑战,包括学习时空信息、处理高维数据、抑制模糊预测以及隐式建模物理特性。在KTH人体动作和Moving-MNIST任务上的大量实验结果表明,尽管模型规模显著更小,但我们的模型在定量与定性评估中均优于主流视频预测技术。最后,我们通过提出一种联合估计HR-VQVAE与ST-PixelCNN参数的新型训练方法进一步提升了S-HR-VQVAE的性能。