Face aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Pluralistic Aging Diffusion Autoencoder (PADA) to enhance the diversity of aging patterns. First, we employ diffusion models to generate diverse low-level aging details via a sequential denoising reverse process. Second, we present Probabilistic Aging Embedding (PAE) to capture diverse high-level aging patterns, which represents age information as probabilistic distributions in the common CLIP latent space. A text-guided KL-divergence loss is designed to guide this learning. Our method can achieve pluralistic face aging conditioned on open-world aging texts and arbitrary unseen face images. Qualitative and quantitative experiments demonstrate that our method can generate more diverse and high-quality plausible aging results.
翻译:人脸衰老是一个病态问题,因为给定输入可能对应多种合理的衰老模式。现有方法通常仅产生单一确定性估计。本文提出一种新颖的CLIP驱动的多元衰老扩散自编码器(PADA),以增强衰老模式的多样性。首先,我们采用扩散模型通过逐步去噪反向过程生成多样化的底层衰老细节。其次,我们提出概率衰老嵌入(PAE)来捕获多样化的高层衰老模式,将年龄信息表示为共享CLIP潜在空间中的概率分布。我们设计了一种文本引导的KL散度损失来指导该学习过程。我们的方法能够在开放世界衰老文本和任意未见人脸图像条件下实现多元人脸衰老。定性与定量实验表明,我们的方法可生成更丰富多样且高质量的可信衰老结果。