This paper proposes a transformer-based learned image compression system. It is capable of achieving variable-rate compression with a single model while supporting the region-of-interest (ROI) functionality. Inspired by prompt tuning, we introduce prompt generation networks to condition the transformer-based autoencoder of compression. Our prompt generation networks generate content-adaptive tokens according to the input image, an ROI mask, and a rate parameter. The separation of the ROI mask and the rate parameter allows an intuitive way to achieve variable-rate and ROI coding simultaneously. Extensive experiments validate the effectiveness of our proposed method and confirm its superiority over the other competing methods.
翻译:本文提出了一种基于Transformer的学习型图像压缩系统。该系统能够通过单一模型实现变速率压缩,同时支持感兴趣区域(ROI)功能。受提示学习启发性,我们引入提示生成网络来调节压缩任务中基于Transformer的自编码器。该提示生成网络可根据输入图像、ROI掩码及速率参数生成内容自适应的提示令牌。通过分离ROI掩码与速率参数,可直观地同时实现变速率编码与ROI编码。大量实验验证了所提方法的有效性,并证实其在同类方法中的优越性能。