The interest of the machine learning community in image synthesis has grown significantly in recent years, with the introduction of a wide range of deep generative models and means for training them. In this work, we propose a general model-agnostic technique for improving the image quality and the distribution fidelity of generated images obtained by any generative model. Our method, termed BIGRoC (Boosting Image Generation via a Robust Classifier), is based on a post-processing procedure via the guidance of a given robust classifier and without a need for additional training of the generative model. Given a synthesized image, we propose to update it through projected gradient steps over the robust classifier to refine its recognition. We demonstrate this post-processing algorithm on various image synthesis methods and show a significant quantitative and qualitative improvement on CIFAR-10 and ImageNet. Surprisingly, although BIGRoC is the first model agnostic among refinement approaches and requires much less information, it outperforms competitive methods. Specifically, BIGRoC improves the image synthesis best performing diffusion model on ImageNet 128x128 by 14.81%, attaining an FID score of 2.53, and on 256x256 by 7.87%, achieving an FID of 3.63. Moreover, we conduct an opinion survey, according to which humans significantly prefer our method's outputs.
翻译:近年来,随着多种深度生成模型及其训练方法的涌现,机器学习社区对图像合成的兴趣显著增长。本文提出一种通用的模型无关技术,用于改善任意生成模型所产出图像的生成质量与分布保真度。我们的方法名为BIGRoC(通过鲁棒分类器提升图像生成),基于给定鲁棒分类器的引导进行后处理,无需额外训练生成模型。针对合成图像,我们提出通过在鲁棒分类器上执行投影梯度步骤来更新图像,从而优化其识别效果。我们在多种图像合成方法上验证了这一后处理算法,并在CIFAR-10和ImageNet数据集上展现了显著的定量与定性提升。令人惊讶的是,尽管BIGRoC是精炼方法中首个模型无关技术且所需信息更少,其性能仍优于竞争方法。具体而言,在ImageNet 128×128分辨率下,BIGRoC将当前最优扩散模型的图像合成性能提升了14.81%,FID分数达到2.53;在256×256分辨率下提升7.87%,FID为3.63。此外,我们进行了意见调查,结果显示人类观察者显著更青睐我们方法的生成结果。