One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despite this potential, a perceptual yet practical image codec is yet to be proposed. In this work, we aim to close this gap. We conduct a comprehensive study of the key modeling choices that govern the design of a practical learned image codec, jointly optimized for perceptual quality and runtime -- including within the ablations several novel techniques. We then perform performance-aware neural architecture search over millions of backbone configurations to identify models that achieve the target on-device runtime while maximizing compression performance as captured by perceptual metrics. We combine the various optimizations to construct a new codec that achieves a significantly improved tradeoff between speed and perceptual quality. Based on rigorous subjective user studies, it provides 2.3-3x bitrate savings against AV1, AV2, VVC, ECM and JPEG-AI, and 20-40% bitrate savings against the best learned codec alternatives. At the same time, on an iPhone 17 Pro Max, it encodes 12MP images as fast as 230ms, and decodes them in 150ms -- faster than most top ML-based codecs run on a V100 GPU.
翻译:与传统硬编码图像编码器相比,学习型编码器的一个关键区别在于其能够直接针对人类视觉系统进行优化。尽管具备这一潜力,目前尚未有兼具感知质量与实用性的图像编码器被提出。本研究旨在填补这一空白。我们系统性地研究了决定实用型学习图像编码器设计的关键建模选择,包括多项创新技术的消融实验,并联合优化了感知质量与运行效率。随后,我们通过性能感知的神经架构搜索,在数百万个主干网络配置中识别出既能达到目标设备运行速度又能最大化感知指标压缩性能的模型。通过融合多种优化策略,我们构建的新型编码器在速度与感知质量之间实现了显著更优的权衡。基于严格的主观用户测试,该编码器相比AV1、AV2、VVC、ECM和JPEG-AI可节省2.3-3倍码率,相比最优的学习型编码器可节省20-40%码率。同时,在iPhone 17 Pro Max上,编码1200万像素图像仅需230毫秒,解码仅需150毫秒——比多数基于机器学习的主流编码器在V100 GPU上的运行速度更快。