The design of a neural image compression network is governed by how well the entropy model matches the true distribution of the latent code. Apart from the model capacity, this ability is indirectly under the effect of how close the relaxed quantization is to the actual hard quantization. Optimizing the parameters of a rate-distortion variational autoencoder (R-D VAE) is ruled by this approximated quantization scheme. In this paper, we propose a feature-level frequency disentanglement to help the relaxed scalar quantization achieve lower bit rates by guiding the high entropy latent features to include most of the low-frequency texture of the image. In addition, to strengthen the de-correlating power of the transformer-based analysis/synthesis transform, an augmented self-attention score calculation based on the Hadamard product is utilized during both encoding and decoding. Channel-wise autoregressive entropy modeling takes advantage of the proposed frequency separation as it inherently directs high-informational low-frequency channels to the first chunks and conditions the future chunks on it. The proposed network not only outperforms hand-engineered codecs, but also neural network-based codecs built on computation-heavy spatially autoregressive entropy models.
翻译:神经图像压缩网络的设计取决于熵模型与潜在编码真实分布的匹配程度。除模型容量外,这种能力间接受到松弛量化与硬量化实际效果之间接近程度的影响。率失真变分自编码器的参数优化受这种近似量化方案的支配。本文提出一种特征级频率解耦方法,通过引导高熵潜在特征包含图像的大部分低频纹理,帮助松弛标量量化实现更低的比特率。此外,为增强基于Transformer的分析/合成变换的去相关能力,在编码和解码过程中采用基于Hadamard积的增强自注意力分数计算。通道级自回归熵建模充分利用了所提出的频率分离特性,因为该特性天然地将高信息量的低频通道引导至前几个块,并对其余块进行条件约束。所提出的网络不仅优于人工设计的编解码器,还超越了基于计算密集型空间自回归熵模型的神经编解码器。