Image Coding for Machines (ICM) is an image compression technique for image recognition. This technique is essential due to the growing demand for image recognition AI. In this paper, we propose a method for ICM that focuses on encoding and decoding only the edge information of object parts in an image, which we call SA-ICM. This is an Learned Image Compression (LIC) model trained using edge information created by Segment Anything. Our method can be used for image recognition models with various tasks. SA-ICM is also robust to changes in input data, making it effective for a variety of use cases. Additionally, our method provides benefits from a privacy point of view, as it removes human facial information on the encoder's side, thus protecting one's privacy. Furthermore, this LIC model training method can be used to train Neural Representations for Videos (NeRV), which is a video compression model. By training NeRV using edge information created by Segment Anything, it is possible to create a NeRV that is effective for image recognition (SA-NeRV). Experimental results confirm the advantages of SA-ICM, presenting the best performance in image compression for image recognition. We also show that SA-NeRV is superior to ordinary NeRV in video compression for machines.
翻译:面向机器图像编码(ICM)是一种用于图像识别的图像压缩技术。随着图像识别人工智能需求的不断增长,该技术至关重要。本文提出了一种名为SA-ICM的ICM方法,其核心思想是仅对图像中物体部分的边缘信息进行编码和解码。该模型是一种基于Segment Anything生成的边缘信息进行训练的 learned image compression(LIC)模型。我们的方法可适用于多种任务的图像识别模型。SA-ICM对输入数据的变化具有鲁棒性,适用于多种应用场景。此外,我们的方法在隐私保护方面具有优势,因为它能在编码端去除人脸信息,从而保护个人隐私。同时,该LIC模型训练方法还可用于训练视频压缩模型Neural Representations for Videos(NeRV)。通过使用Segment Anything生成的边缘信息训练NeRV,可以构建对图像识别有效的NeRV(SA-NeRV)。实验结果表明,SA-ICM在面向图像识别的图像压缩中具有最佳性能,且SA-NeRV在面向机器视频压缩中优于普通NeRV。