Synthesizing face images from monochrome sketches is one of the most fundamental tasks in the field of image-to-image translation. However, it is still challenging to (1)~make models learn the high-dimensional face features such as geometry and color, and (2)~take into account the characteristics of input sketches. Existing methods often use sketches as indirect inputs (or as auxiliary inputs) to guide the models, resulting in the loss of sketch features or the alteration of geometry information. In this paper, we introduce a Sketch-Guided Latent Diffusion Model (SGLDM), an LDM-based network architect trained on the paired sketch-face dataset. We apply a Multi-Auto-Encoder (AE) to encode the different input sketches from different regions of a face from pixel space to a feature map in latent space, which enables us to reduce the dimension of the sketch input while preserving the geometry-related information of local face details. We build a sketch-face paired dataset based on the existing method that extracts the edge map from an image. We then introduce a Stochastic Region Abstraction (SRA), an approach to augment our dataset to improve the robustness of SGLDM to handle sketch input with arbitrary abstraction. The evaluation study shows that SGLDM can synthesize high-quality face images with different expressions, facial accessories, and hairstyles from various sketches with different abstraction levels.
翻译:从单色草图中合成人脸图像是图像到图像翻译领域中最基础的任务之一。然而,让模型学习几何、颜色等高维人脸特征,同时考虑输入草图的特性仍然具有挑战性。现有方法通常将草图作为间接输入(或辅助输入)引导模型,导致草图特征丢失或几何信息改变。本文提出了一种草图引导潜在扩散模型(SGLDM),这是一种基于潜在扩散模型(LDM)的网络架构,在配对的人脸草图-照片数据集上训练。我们采用多自编码器(Multi-Auto-Encoder)将来自人脸不同区域的输入草图从像素空间编码到潜在空间的特征图,从而在保留局部人脸细节几何相关信息的同时降低草图输入的维度。基于现有从图像提取边缘图的方法,我们构建了草图-人脸配对数据集。随后引入随机区域抽象(SRA)方法增强数据集,以提升SGLDM处理任意抽象程度草图输入的鲁棒性。评估研究表明,SGLDM能够从不同抽象级别的草图中合成具有不同表情、面部配饰及发型的高质量人脸图像。