We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD, considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion compared to PLF-FMD and further makes it more difficult to recover from errors made in the earlier output frames. Although the choice of PLF decisively affects reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be {\em near universal} and can generate close to optimal reconstructions for either choice of PLF at the decoder. We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST and KTH datasets.
翻译:我们研究了在输出同时受均方误差(MSE)失真损失和感知损失(以提升真实感为目标)约束下的因果、低延迟、顺序视频压缩问题。受已有方法启发,我们考虑两种不同的感知损失函数(PLF)。第一种PLF-JD关注所有当前帧之前视频帧的联合分布(JD),第二种PLF-FMD则关注源与重建之间的逐帧边际分布(FMD)。通过信息论分析与基于深度学习的实验,我们证明PLF的选择会对重建质量产生显著影响,尤其是在低比特率条件下。具体而言,尽管基于PLF-JD的重建能更好地保留帧间时间相关性,但与PLF-FMD相比,其失真惩罚更大,且更难从前序输出帧的错误中恢复。尽管PLF的选择显著影响重建质量,我们进一步证明编码过程中无需固定采用特定PLF,且PLF的选择可移交给解码器。特别地,通过训练系统最小化MSE(无需任何PLF)生成的编码表示具有“近通用性”,能在解码端针对任选PLF实现近乎最优的重建效果。我们通过(单次)信息论分析、高斯-马尔可夫源模型的率-失真-感知权衡详细研究,以及基于移动MNIST和KTH数据集的深度学习实验,验证了上述结论。