Recent progress in generative compression technology has significantly improved the perceptual quality of compressed data. However, these advancements primarily focus on producing high-frequency details, often overlooking the ability of generative models to capture the prior distribution of image content, thus impeding further bitrate reduction in extreme compression scenarios (<0.05 bpp). Motivated by the capabilities of predictive language models for lossless compression, this paper introduces a novel Unified Image Generation-Compression (UIGC) paradigm, merging the processes of generation and compression. A key feature of the UIGC framework is the adoption of vector-quantized (VQ) image models for tokenization, alongside a multi-stage transformer designed to exploit spatial contextual information for modeling the prior distribution. As such, the dual-purpose framework effectively utilizes the learned prior for entropy estimation and assists in the regeneration of lost tokens. Extensive experiments demonstrate the superiority of the proposed UIGC framework over existing codecs in perceptual quality and human perception, particularly in ultra-low bitrate scenarios (<=0.03 bpp), pioneering a new direction in generative compression.
翻译:近期生成式压缩技术的进展显著提升了压缩数据的感知质量。然而,这些进展主要集中在高频细节生成上,往往忽略了生成模型对图像内容先验分布的捕获能力,从而阻碍了在极端压缩场景(<0.05 bpp)中进一步降低码率。受预测式语言模型在无损压缩中能力的启发,本文提出一种新型统一图像生成-压缩(UIGC)范式,将生成与压缩过程相融合。该框架的核心在于采用向量量化(VQ)图像模型进行词元化,同时引入多阶段Transformer以利用空间上下文信息建模先验分布。由此,这一双用途框架能够有效利用所学习的先验进行熵估计,并辅助丢失词元的再生。大量实验表明,所提出的UIGC框架在感知质量与人类感知上优于现有编解码器,尤其是在超低码率场景(<=0.03 bpp)中,开创了生成式压缩的新方向。