Image compression aims to reduce the information redundancy in images. Most existing neural image compression methods rely on side information from hyperprior or context models to eliminate spatial redundancy, but rarely address the channel redundancy. Inspired by the mask sampling modeling in recent self-supervised learning methods for natural language processing and high-level vision, we propose a novel pretraining strategy for neural image compression. Specifically, Cube Mask Sampling Module (CMSM) is proposed to apply both spatial and channel mask sampling modeling to image compression in the pre-training stage. Moreover, to further reduce channel redundancy, we propose the Learnable Channel Mask Module (LCMM) and the Learnable Channel Completion Module (LCCM). Our plug-and-play CMSM, LCMM, LCCM modules can apply to both CNN-based and Transformer-based architectures, significantly reduce the computational cost, and improve the quality of images. Experiments on the public Kodak and Tecnick datasets demonstrate that our method achieves competitive performance with lower computational complexity compared to state-of-the-art image compression methods.
翻译:图像压缩旨在减少图像中的信息冗余。现有的大多数神经图像压缩方法依赖超先验或上下文模型提供的边信息来消除空间冗余,但很少处理通道冗余。受自然语言处理和高层视觉中近期自监督学习方法中掩码采样建模的启发,我们为神经图像压缩提出了一种新颖的预训练策略。具体而言,我们提出了立方体掩码采样模块(CMSM),在预训练阶段对图像压缩同时进行空间和通道掩码采样建模。此外,为进一步减少通道冗余,我们提出了可学习通道掩码模块(LCMM)和可学习通道补全模块(LCCM)。我们的即插即用模块CMSM、LCMM和LCCM可适用于基于CNN和基于Transformer的架构,显著降低计算成本并提升图像质量。在公开的Kodak和Tecnick数据集上的实验表明,与最先进的图像压缩方法相比,我们的方法在实现竞争性能的同时降低了计算复杂度。