We offer a method for one-shot mask-guided image synthesis that allows controlling manipulations of a single image by inverting a quasi-robust classifier equipped with strong regularizers. Our proposed method, entitled MAGIC, leverages structured gradients from a pre-trained quasi-robust classifier to better preserve the input semantics while preserving its classification accuracy, thereby guaranteeing credibility in the synthesis. Unlike current methods that use complex primitives to supervise the process or use attention maps as a weak supervisory signal, MAGIC aggregates gradients over the input, driven by a guide binary mask that enforces a strong, spatial prior. MAGIC implements a series of manipulations with a single framework achieving shape and location control, intense non-rigid shape deformations, and copy/move operations in the presence of repeating objects and gives users firm control over the synthesis by requiring to simply specify binary guide masks. Our study and findings are supported by various qualitative comparisons with the state-of-the-art on the same images sampled from ImageNet and quantitative analysis using machine perception along with a user survey of 100+ participants that endorse our synthesis quality. Project page at https://mozhdehrouhsedaghat.github.io/magic.html. Code is available at https://github.com/mozhdehrouhsedaghat/magic
翻译:我们提出了一种一次性掩码引导图像合成方法,通过逆变换一个配备强正则化器的高斯稳健分类器,实现对单幅图像的可控操作。本方法名为MAGIC,利用预训练准鲁棒分类器的结构化梯度,在保持分类精度的同时更好地保留输入语义,从而保证合成结果的可信度。与当前使用复杂基元监督过程或依赖注意力图作为弱监督信号的方法不同,MAGIC通过引导二值掩码施加强空间先验,将输入上的梯度进行聚合。该方法通过单一框架实现一系列操作:控制形状与位置、实现强非刚性形变、针对重复目标的复制/移动操作,用户仅需指定简单的二值引导掩码即可获得对合成过程的精确控制。我们的研究通过多项定性比较(与当前最优方法在ImageNet采样图像上的对比)、基于机器感知的定量分析以及包含100+参与者的用户调查验证了合成质量。项目页面:https://mozhdehrouhsedaghat.github.io/magic.html。代码地址:https://github.com/mozhdehrouhsedaghat/magic