By optimizing the rate-distortion-realism trade-off, generative compression approaches produce detailed, realistic images, even at low bit rates, instead of the blurry reconstructions produced by rate-distortion optimized models. However, previous methods do not explicitly control how much detail is synthesized, which results in a common criticism of these methods: users might be worried that a misleading reconstruction far from the input image is generated. In this work, we alleviate these concerns by training a decoder that can bridge the two regimes and navigate the distortion-realism trade-off. From a single compressed representation, the receiver can decide to either reconstruct a low mean squared error reconstruction that is close to the input, a realistic reconstruction with high perceptual quality, or anything in between. With our method, we set a new state-of-the-art in distortion-realism, pushing the frontier of achievable distortion-realism pairs, i.e., our method achieves better distortions at high realism and better realism at low distortion than ever before.
翻译:通过优化率-失真-现实主义权衡,生成式压缩方法即使在低比特率下也能生成细节丰富、逼真的图像,而非由率-失真优化模型产生的模糊重建图像。然而,以往方法并未明确控制合成细节的数量,这导致了对这些方法的常见批评:用户可能担心生成远离输入图像的误导性重建。在本工作中,我们通过训练一个能衔接两种机制并驾驭失真-现实主义权衡的解码器来缓解这些担忧。从单一压缩表示出发,接收方可决定重建出与输入接近的低均方误差重建、高感知质量的逼真重建,或介于两者之间的任意结果。借助我们的方法,我们在失真-现实主义方面设立了新的最优水平,将可实现的失真-现实主义对的前沿推进一步,即:相较于以往,我们的方法在高现实主义下实现了更优的失真,在低失真下实现了更优的现实主义。