3D-aware portrait editing has a wide range of applications in multiple fields. However, current approaches are limited due that they can only perform mask-guided or text-based editing. Even by fusing the two procedures into a model, the editing quality and stability cannot be ensured. To address this limitation, we propose \textbf{MaTe3D}: mask-guided text-based 3D-aware portrait editing. In this framework, first, we introduce a new SDF-based 3D generator which learns local and global representations with proposed SDF and density consistency losses. This enhances masked-based editing in local areas; second, we present a novel distillation strategy: Conditional Distillation on Geometry and Texture (CDGT). Compared to exiting distillation strategies, it mitigates visual ambiguity and avoids mismatch between texture and geometry, thereby producing stable texture and convincing geometry while editing. Additionally, we create the CatMask-HQ dataset, a large-scale high-resolution cat face annotation for exploration of model generalization and expansion. We perform expensive experiments on both the FFHQ and CatMask-HQ datasets to demonstrate the editing quality and stability of the proposed method. Our method faithfully generates a 3D-aware edited face image based on a modified mask and a text prompt. Our code and models will be publicly released.
翻译:三维感知肖像编辑在多个领域具有广泛应用。然而,现有方法受限于仅能执行掩码引导或文本驱动的编辑。即便将两种流程融合至同一模型,其编辑质量与稳定性仍无法保证。为解决此问题,我们提出\textbf{MaTe3D}:基于掩码引导的文本驱动三维感知肖像编辑框架。在该框架中,首先引入一种新型基于SDF的三维生成器,通过所提出的SDF与密度一致性损失学习局部与全局表征,从而增强局部区域的掩码式编辑能力;其次,提出新型蒸馏策略——几何与纹理条件蒸馏(CDGT)。相较于现有蒸馏策略,该方法可缓解视觉模糊性并避免纹理与几何的失配,进而在编辑过程中生成稳定的纹理与可信的几何结构。此外,我们构建了CatMask-HQ数据集——大规模高分辨率猫脸标注数据集,用于探索模型泛化性与扩展性。通过在FFHQ与CatMask-HQ数据集上的大量实验,验证了所提方法的编辑质量与稳定性。我们的方法能够基于修改后的掩码与文本提示,忠实生成三维感知编辑后的面部图像。代码与模型将公开发布。