The images produced by diffusion models can attain excellent perceptual quality. However, it is challenging for diffusion models to guarantee distortion, hence the integration of diffusion models and image compression models still needs more comprehensive explorations. This paper presents a diffusion-based image compression method that employs a privileged end-to-end decoder model as correction, which achieves better perceptual quality while guaranteeing the distortion to an extent. We build a diffusion model and design a novel paradigm that combines the diffusion model and an end-to-end decoder, and the latter is responsible for transmitting the privileged information extracted at the encoder side. Specifically, we theoretically analyze the reconstruction process of the diffusion models at the encoder side with the original images being visible. Based on the analysis, we introduce an end-to-end convolutional decoder to provide a better approximation of the score function $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at the encoder side and effectively transmit the combination. Experiments demonstrate the superiority of our method in both distortion and perception compared with previous perceptual compression methods.
翻译:扩散模型生成的图像能够达到优异的感知质量。然而,扩散模型难以保证失真性能,因此扩散模型与图像压缩模型的结合仍需更全面的探索。本文提出一种基于扩散的图像压缩方法,该方法采用特权端到端解码器模型作为校正,在保证一定失真程度的同时实现了更优的感知质量。我们构建了一个扩散模型,并设计了一种结合扩散模型与端到端解码器的新范式,后者负责传输在编码端提取的特权信息。具体而言,我们从理论上分析了原始图像可见情况下编码端扩散模型的重建过程。基于该分析,我们引入一个端到端卷积解码器,以在编码端对得分函数 $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ 提供更优的近似,并有效传输组合信息。实验结果表明,与先前的感知压缩方法相比,本方法在失真和感知指标上均具有优越性。