Image coding for machines (ICM) aims to compress images to support downstream AI analysis instead of human perception. For ICM, developing a unified codec to reduce information redundancy while empowering the compressed features to support various vision tasks is very important, which inevitably faces two core challenges: 1) How should the compression strategy be adjusted based on the downstream tasks? 2) How to well adapt the compressed features to different downstream tasks? Inspired by recent advances in transferring large-scale pre-trained models to downstream tasks via prompting, in this work, we explore a new ICM framework, termed Prompt-ICM. To address both challenges by carefully learning task-driven prompts to coordinate well the compression process and downstream analysis. Specifically, our method is composed of two core designs: a) compression prompts, which are implemented as importance maps predicted by an information selector, and used to achieve different content-weighted bit allocations during compression according to different downstream tasks; b) task-adaptive prompts, which are instantiated as a few learnable parameters specifically for tuning compressed features for the specific intelligent task. Extensive experiments demonstrate that with a single feature codec and a few extra parameters, our proposed framework could efficiently support different kinds of intelligent tasks with much higher coding efficiency.
翻译:机器图像编码旨在压缩图像以支持下游AI分析,而非人类感知。对于机器图像编码而言,开发一种统一编解码器,在减少信息冗余的同时使压缩特征能够支持多种视觉任务至关重要,这不可避免地面临两个核心挑战:1)如何根据下游任务调整压缩策略?2)如何使压缩特征良好适配不同的下游任务?受近期通过提示将大规模预训练模型迁移至下游任务的研究启发,本文探索了一种新的机器图像编码框架,称为Prompt-ICM。通过精心学习任务驱动型提示来协调压缩过程与下游分析,以解决上述两个挑战。具体而言,我们的方法包含两个核心设计:a)压缩提示——由信息选择器预测的重要性图实现,用于根据不同下游任务在压缩过程中实现不同的内容加权比特分配;b)任务自适应提示——实例化为少量可学习参数,专门用于针对特定智能任务微调压缩特征。大量实验表明,仅需单一特征编解码器和少量额外参数,我们的框架即可高效支持多种智能化任务,并显著提升编码效率。