Diffusion models have shown promising results on single-image super-resolution and other image- to-image translation tasks. Despite this success, they have not outperformed state-of-the-art GAN models on the more challenging blind super-resolution task, where the input images are out of distribution, with unknown degradations. This paper introduces SR3+, a diffusion-based model for blind super-resolution, establishing a new state-of-the-art. To this end, we advocate self-supervised training with a combination of composite, parameterized degradations for self-supervised training, and noise-conditioing augmentation during training and testing. With these innovations, a large-scale convolutional architecture, and large-scale datasets, SR3+ greatly outperforms SR3. It outperforms Real-ESRGAN when trained on the same data, with a DRealSR FID score of 36.82 vs. 37.22, which further improves to FID of 32.37 with larger models, and further still with larger training sets.
翻译:扩散模型在单图像超分辨率及其他图像到图像翻译任务中已展现出 promising 的结果。尽管取得了这一成功,但在更具挑战性的盲超分辨率任务中,它们尚未超越最先进的GAN模型——该类任务中的输入图像分布外,且存在未知退化。本文提出SR3+,一种基于扩散的盲超分辨率模型,建立了新的最先进水平。为此,我们倡导采用组合式参数化退化的自监督训练,并在训练和测试过程中引入噪声条件增强。凭借这些创新、大规模卷积架构及大规模数据集,SR3+大幅超越了SR3。在相同数据上训练时,它优于Real-ESRGAN,DRealSR FID得分为36.82对比37.22,通过采用更大模型可进一步提升至FID 32.37,而使用更大训练集时表现更优。