We introduce an approach for 3D head avatar generation and editing with multi-modal conditioning based on a 3D Generative Adversarial Network (GAN) and a Latent Diffusion Model (LDM). 3D GANs can generate high-quality head avatars given a single or no condition. However, it is challenging to generate samples that adhere to multiple conditions of different modalities. On the other hand, LDMs excel at learning complex conditional distributions. To this end, we propose to exploit the conditioning capabilities of LDMs to enable multi-modal control over the latent space of a pre-trained 3D GAN. Our method can generate and edit 3D head avatars given a mixture of control signals such as RGB input, segmentation masks, and global attributes. This provides better control over the generation and editing of synthetic avatars both globally and locally. Experiments show that our proposed approach outperforms a solely GAN-based approach both qualitatively and quantitatively on generation and editing tasks. To the best of our knowledge, our approach is the first to introduce multi-modal conditioning to 3D avatar generation and editing. \\href{avatarmmc-sig24.github.io}{Project Page}
翻译:我们提出了一种基于3D生成对抗网络(GAN)和潜在扩散模型(LDM)的多模态条件驱动的3D头部虚拟人生成与编辑方法。3D GAN能够在单一条件或零条件下生成高质量的头部虚拟人,但在生成符合不同模态多重条件的样本时面临挑战。另一方面,LDM擅长学习复杂的条件分布。为此,我们提出利用LDM的条件控制能力,实现对预训练3D GAN潜在空间的多模态操控。我们的方法能够根据RGB输入、分割掩码和全局属性等混合控制信号,生成并编辑3D头部虚拟人。这为合成虚拟人的全局与局部生成及编辑提供了更优的控制能力。实验表明,在生成与编辑任务中,本方法在定性和定量指标上均优于纯GAN方法。据我们所知,本方法是首个将多模态条件引入3D虚拟人生成与编辑的研究。\href{avatarmmc-sig24.github.io}{项目页面}