Recent years have witnessed the dramatic growth of Internet video traffic, where the video bitstreams are often compressed and delivered in low quality to fit the streamer's uplink bandwidth. To alleviate the quality degradation, it comes the rise of Neural-enhanced Video Streaming (NVS), which shows great prospects for recovering low-quality videos by mostly deploying neural super-resolution (SR) on the media server. Despite its benefit, we reveal that current mainstream works with SR enhancement have not achieved the desired rate-distortion trade-off between bitrate saving and quality restoration, due to: (1) overemphasizing the enhancement on the decoder side while omitting the co-design of encoder, (2) limited generative capacity to recover high-fidelity perceptual details, and (3) optimizing the compression-and-restoration pipeline from the resolution perspective solely, without considering color bit-depth. Aiming at overcoming these limitations, we are the first to conduct an encoder-decoder (i.e., codec) synergy by leveraging the inherent visual-generative property of diffusion models. Specifically, we present the Codec-aware Diffusion Modeling (CaDM), a novel NVS paradigm to significantly reduce streaming delivery bitrates while holding pretty higher restoration capacity over existing methods. First, CaDM improves the encoder's compression efficiency by simultaneously reducing resolution and color bit-depth of video frames. Second, CaDM empowers the decoder with high-quality enhancement by making the denoising diffusion restoration aware of encoder's resolution-color conditions. Evaluation on public cloud services with OpenMMLab benchmarks shows that CaDM effectively saves up to 5.12 - 21.44 times bitrates based on common video standards and achieves much better recovery quality (e.g., FID of 0.61) over state-of-the-art neural-enhancing methods.
翻译:摘要:近年来,互联网视频流量急剧增长,视频比特流通常被压缩并以低质量传输,以适应流媒体上行带宽限制。为缓解质量下降,神经增强视频流(NVS)应运而生,其通过在媒体服务器端部署神经超分辨率(SR)技术,在恢复低质量视频方面展现出巨大前景。然而,尽管具有优势,我们揭示当前基于SR增强的主流方法并未在比特率节省与质量恢复之间实现理想的率失真权衡,原因包括:(1)过度强调解码端增强而忽略编码器的协同设计;(2)生成能力有限,难以恢复高保真感知细节;(3)仅从分辨率角度优化压缩-恢复管线,而未考虑色彩位深度。为克服这些局限,我们首次利用扩散模型固有的视觉生成特性,实现编码器-解码器(即编解码器)协同。具体而言,我们提出编解码感知扩散建模(CaDM),一种新型NVS范式,在显著降低流媒体传输比特率的同时,相比现有方法保持更高的恢复能力。首先,CaDM通过同步降低视频帧的分辨率和色彩位深度,提升编码器压缩效率。其次,CaDM通过使去噪扩散恢复过程感知编码器的分辨率-色彩条件,赋予解码器高质量增强能力。基于OpenMMLab基准在公共云服务上的评估表明,CaDM基于常见视频标准可有效节省5.12至21.44倍比特率,并在恢复质量上(例如FID达0.61)显著优于最先进的神经增强方法。