Given an abstract, deformed, ordinary sketch from untrained amateurs like you and me, this paper turns it into a photorealistic image - just like those shown in Fig. 1(a), all non-cherry-picked. We differ significantly from prior art in that we do not dictate an edgemap-like sketch to start with, but aim to work with abstract free-hand human sketches. In doing so, we essentially democratise the sketch-to-photo pipeline, "picturing" a sketch regardless of how good you sketch. Our contribution at the outset is a decoupled encoder-decoder training paradigm, where the decoder is a StyleGAN trained on photos only. This importantly ensures that generated results are always photorealistic. The rest is then all centred around how best to deal with the abstraction gap between sketch and photo. For that, we propose an autoregressive sketch mapper trained on sketch-photo pairs that maps a sketch to the StyleGAN latent space. We further introduce specific designs to tackle the abstract nature of human sketches, including a fine-grained discriminative loss on the back of a trained sketch-photo retrieval model, and a partial-aware sketch augmentation strategy. Finally, we showcase a few downstream tasks our generation model enables, amongst them is showing how fine-grained sketch-based image retrieval, a well-studied problem in the sketch community, can be reduced to an image (generated) to image retrieval task, surpassing state-of-the-arts. We put forward generated results in the supplementary for everyone to scrutinise.
翻译:本文提出将抽象、变形、随意的手绘草图(如同你我这样的非专业业余爱好者所画)转化为逼真的照片级图像——如图1(a)所示,所有结果均非精心挑选。与现有方法显著不同的是,我们不要求输入类边缘图的规范草图,而是致力于处理抽象的徒手人绘草图。通过此举,我们实质上实现了草图到照片管线的民主化,无论用户绘图水平如何,都能将草图“照片化”。本文的核心贡献在于提出解耦的编码器-解码器训练范式,其中解码器是仅在照片上训练的StyleGAN。这一设计确保生成结果始终具有照片级真实感。其余工作围绕如何弥合草图与照片之间的抽象差距展开。为此,我们提出自回归草图映射器,该模型在成对的草图-照片数据上训练,可将草图映射到StyleGAN隐空间。针对人类草图的抽象特性,我们进一步引入特定设计:基于预训练的草图-照片检索模型的细粒度判别损失,以及部分感知的草图增强策略。最后,我们展示了生成模型支持的几个下游任务,其中包括证明:草图领域中被广泛研究的细粒度草图检索问题,可简化为图像(生成图像)到图像的检索任务,且超越现有方法。生成结果详见补充材料,以供各位审视。