Text-guided domain adaptation and generation of 3D-aware portraits find many applications in various fields. However, due to the lack of training data and the challenges in handling the high variety of geometry and appearance, the existing methods for these tasks suffer from issues like inflexibility, instability, and low fidelity. In this paper, we propose a novel framework DiffusionGAN3D, which boosts text-guided 3D domain adaptation and generation by combining 3D GANs and diffusion priors. Specifically, we integrate the pre-trained 3D generative models (e.g., EG3D) and text-to-image diffusion models. The former provides a strong foundation for stable and high-quality avatar generation from text. And the diffusion models in turn offer powerful priors and guide the 3D generator finetuning with informative direction to achieve flexible and efficient text-guided domain adaptation. To enhance the diversity in domain adaptation and the generation capability in text-to-avatar, we introduce the relative distance loss and case-specific learnable triplane respectively. Besides, we design a progressive texture refinement module to improve the texture quality for both tasks above. Extensive experiments demonstrate that the proposed framework achieves excellent results in both domain adaptation and text-to-avatar tasks, outperforming existing methods in terms of generation quality and efficiency. The project homepage is at https://younglbw.github.io/DiffusionGAN3D-homepage/.
翻译:文本引导的域自适应与三维感知人像生成在众多领域具有广泛应用前景。然而,由于训练数据匮乏以及处理几何与外观高多样性的挑战,现有方法在灵活性、稳定性和保真度方面仍存在不足。本文提出新型框架DiffusionGAN3D,通过融合3D生成对抗网络与扩散先验,显著提升文本引导的三维域自适应与生成性能。具体而言,我们整合了预训练的三维生成模型(如EG3D)与文本到图像扩散模型:前者为基于文本的高质量稳定虚拟人像生成奠定坚实基础;后者则提供强大先验信息,通过有效方向引导三维生成器的微调,实现灵活高效的文本引导域自适应。为增强域自适应的多样性及文本到虚拟人像的生成能力,我们分别引入相对距离损失与任务特定可学习三平面。此外,设计渐进式纹理精炼模块以改善两个任务的纹理质量。大量实验表明,所提框架在域自适应与文本到虚拟人像两大任务中均取得优异效果,在生成质量与效率上超越现有方法。项目主页见https://younglbw.github.io/DiffusionGAN3D-homepage/。