We propose sandwiched video compression -- a video compression system that wraps neural networks around a standard video codec. The sandwich framework consists of a neural pre- and post-processor with a standard video codec between them. The networks are trained jointly to optimize a rate-distortion loss function with the goal of significantly improving over the standard codec in various compression scenarios. End-to-end training in this setting requires a differentiable proxy for the standard video codec, which incorporates temporal processing with motion compensation, inter/intra mode decisions, and in-loop filtering. We propose differentiable approximations to key video codec components and demonstrate that the neural codes of the sandwich lead to significantly better rate-distortion performance compared to compressing the original frames of the input video in two important scenarios. When transporting high-resolution video via low-resolution HEVC, the sandwich system obtains 6.5 dB improvements over standard HEVC. More importantly, using the well-known perceptual similarity metric, LPIPS, we observe $~30 \%$ improvements in rate at the same quality over HEVC. Last but not least we show that pre- and post-processors formed by very modestly-parameterized, light-weight networks can closely approximate these results.
翻译:我们提出夹层视频压缩——一种通过将神经网络环绕标准视频编解码器构建的视频压缩系统。该夹层框架包含神经预处理与后处理模块,中间嵌入标准视频编解码器。网络通过联合训练优化率失真损失函数,旨在多种压缩场景下显著超越标准编解码器性能。在此设置中进行端到端训练需要标准视频编解码器的可微分代理,该代理需整合包含运动补偿、帧内/帧间模式决策及环路滤波的时间处理功能。我们对关键视频编解码组件提出可微分近似方案,并证明夹层中的神经编码在两种重要场景下能显著提升率失真性能:当通过低分辨率HEVC传输高分辨率视频时,夹层系统相较标准HEVC获得6.5 dB增益;更重要的是,采用公认的感知相似度度量LPIPS,在相同质量下相较HEVC可降低约30%码率。最后我们证实,采用参数极简的轻量级网络构建前后处理模块即可高度逼近上述结果。