Recent advances in face manipulation using StyleGAN have produced impressive results. However, StyleGAN is inherently limited to cropped aligned faces at a fixed image resolution it is pre-trained on. In this paper, we propose a simple and effective solution to this limitation by using dilated convolutions to rescale the receptive fields of shallow layers in StyleGAN, without altering any model parameters. This allows fixed-size small features at shallow layers to be extended into larger ones that can accommodate variable resolutions, making them more robust in characterizing unaligned faces. To enable real face inversion and manipulation, we introduce a corresponding encoder that provides the first-layer feature of the extended StyleGAN in addition to the latent style code. We validate the effectiveness of our method using unaligned face inputs of various resolutions in a diverse set of face manipulation tasks, including facial attribute editing, super-resolution, sketch/mask-to-face translation, and face toonification.
翻译:近期利用StyleGAN进行人脸操作的研究取得了显著成果,但StyleGAN本质上受限于其预训练时固定的图像分辨率所对应的裁剪对齐人脸。本文提出一种简单有效的解决方案:通过使用扩张卷积调整StyleGAN浅层特征的感受野,无需修改任何模型参数。该方法可将浅层固定尺寸的小尺度特征扩展为能适应可变分辨率的大尺度特征,从而增强其对未对齐人脸的鲁棒表征能力。为实现真实人脸反演与操作,我们引入对应的编码器,该编码器除潜空间风格编码外还能提供扩展后StyleGAN的首层特征。通过多种人脸操作任务(包括面部属性编辑、超分辨率、草图/蒙版到人脸合成、人脸卡通化)中多分辨率未对齐人脸输入验证了方法的有效性。