Generating images from human sketches typically requires dedicated networks trained from scratch. In contrast, the emergence of the pre-trained Vision-Language models (e.g., CLIP) has propelled generative applications based on controlling the output imagery of existing StyleGAN models with text inputs or reference images. Parallelly, our work proposes a framework to control StyleGAN imagery with a single user sketch. In particular, we learn a conditional distribution in the latent space of a pre-trained StyleGAN model via energy-based learning and propose two novel energy functions leveraging CLIP for cross-domain semantic supervision. Once trained, our model can generate multi-modal images semantically aligned with the input sketch. Quantitative evaluations on synthesized datasets have shown that our approach improves significantly from previous methods in the one-shot regime. The superiority of our method is further underscored when experimenting with a wide range of human sketches of diverse styles and poses. Surprisingly, our models outperform the previous baseline regarding both the range of sketch inputs and image qualities despite operating with a stricter setting: with no extra training data and single sketch input.
翻译:从人类手绘草稿生成图像通常需要从头训练专用的网络。相比之下,预训练视觉-语言模型(如CLIP)的出现推动了基于文本输入或参考图像控制现有StyleGAN模型输出图像的生成应用。与此并行,我们的工作提出了一种框架,通过单张用户手绘草稿控制StyleGAN图像生成。具体而言,我们通过基于能量的学习在预训练StyleGAN模型的潜在空间中学习条件分布,并提出了两种利用CLIP进行跨域语义监督的新型能量函数。经过训练后,我们的模型能够生成与输入草稿语义对齐的多模态图像。在合成数据集上的定量评估表明,我们的方法在单样本场景下相较于先前方法有显著提升。当对多种风格和姿态的人类手绘草稿进行实验时,我们方法的优越性进一步凸显。令人惊讶的是,尽管在更严格的设置下(无额外训练数据且仅使用单张草稿输入),我们的模型在草稿输入范围与图像质量两方面均超越了先前基线方法。