While replacing Gaussian decoders with a conditional diffusion model enhances the perceptual quality of reconstructions in neural image compression, their lack of inductive bias for image data restricts their ability to achieve state-of-the-art perceptual levels. To address this limitation, we adopt a non-isotropic diffusion model at the decoder side. This model imposes an inductive bias aimed at distinguishing between frequency contents, thereby facilitating the generation of high-quality images. Moreover, our framework is equipped with a novel entropy model that accurately models the probability distribution of latent representation by exploiting spatio-channel correlations in latent space, while accelerating the entropy decoding step. This channel-wise entropy model leverages both local and global spatial contexts within each channel chunk. The global spatial context is built upon the Transformer, which is specifically designed for image compression tasks. The designed Transformer employs a Laplacian-shaped positional encoding, the learnable parameters of which are adaptively adjusted for each channel cluster. Our experiments demonstrate that our proposed framework yields better perceptual quality compared to cutting-edge generative-based codecs, and the proposed entropy model contributes to notable bitrate savings.
翻译:虽然用条件扩散模型替代高斯解码器能提升神经图像压缩中重建图像的感知质量,但这类模型缺乏针对图像数据的归纳偏置,限制了其达到最先进感知水平的能力。为解决这一局限,我们在解码端采用非各向同性扩散模型。该模型通过施加能区分频率内容的归纳偏置,促进高质量图像的生成。此外,我们的框架配备了新型熵模型,该模型通过挖掘潜在空间中的时空通道相关性,精准建模潜在表示的概率分布,同时加速熵解码步骤。这种通道级熵模型在每个通道块内同时利用局部和全局空间上下文。全局空间上下文基于针对图像压缩任务特制的Transformer构建,该Transformer采用拉普拉斯形位置编码,其可学习参数可针对每个通道簇自适应调节。实验表明,与前沿生成式编解码器相比,我们提出的框架能带来更优的感知质量,且所提熵模型有助于显著节省码率。