Facial Appearance Editing (FAE) aims to modify physical attributes, such as pose, expression and lighting, of human facial images while preserving attributes like identity and background, showing great importance in photograph. In spite of the great progress in this area, current researches generally meet three challenges: low generation fidelity, poor attribute preservation, and inefficient inference. To overcome above challenges, this paper presents DiffFAE, a one-stage and highly-efficient diffusion-based framework tailored for high-fidelity FAE. For high-fidelity query attributes transfer, we adopt Space-sensitive Physical Customization (SPC), which ensures the fidelity and generalization ability by utilizing rendering texture derived from 3D Morphable Model (3DMM). In order to preserve source attributes, we introduce the Region-responsive Semantic Composition (RSC). This module is guided to learn decoupled source-regarding features, thereby better preserving the identity and alleviating artifacts from non-facial attributes such as hair, clothes, and background. We further introduce a consistency regularization for our pipeline to enhance editing controllability by leveraging prior knowledge in the attention matrices of diffusion model. Extensive experiments demonstrate the superiority of DiffFAE over existing methods, achieving state-of-the-art performance in facial appearance editing.
翻译:面部外观编辑(FAE)旨在修改人脸图像的物理属性(如姿态、表情和光照),同时保持身份和背景等属性,在摄影领域具有重要价值。尽管该领域已取得显著进展,现有研究普遍面临三大挑战:生成保真度低、属性保持差、推理效率低。为克服上述挑战,本文提出DiffFAE——一种面向高保真FAE的单阶段高效扩散框架。针对高保真查询属性迁移,我们采用空间敏感物理定制(SPC)模块,通过利用三维可变形模型(3DMM)生成的渲染纹理,确保保真度与泛化能力。为保持源属性,我们引入区域响应语义组合(RSC)模块,该模块通过学习解耦的源属性特征,可更好地保持身份信息并缓解头发、衣物、背景等非面部属性产生的伪影。此外,我们通过利用扩散模型注意力矩阵中的先验知识引入一致性正则化,以增强编辑可控性。大量实验表明,DiffFAE在面部外观编辑任务中显著超越现有方法,达到当前最优水平。