Domain adaptation of 3D portraits has gained more and more attention. However, the transfer mechanism of existing methods is mainly based on vision or language, which ignores the potential of vision-language combined guidance. In this paper, we propose an Image-Text multi-modal framework, namely Image and Text portrait (ITportrait), for 3D portrait domain adaptation. ITportrait relies on a two-stage alternating training strategy. In the first stage, we employ a 3D Artistic Paired Transfer (APT) method for image-guided style transfer. APT constructs paired photo-realistic portraits to obtain accurate artistic poses, which helps ITportrait to achieve high-quality 3D style transfer. In the second stage, we propose a 3D Image-Text Embedding (ITE) approach in the CLIP space. ITE uses a threshold function to self-adaptively control the optimization direction of images or texts in the CLIP space. Comprehensive experiments prove that our ITportrait achieves state-of-the-art (SOTA) results and benefits downstream tasks. All source codes and pre-trained models will be released to the public.
翻译:三维肖像的域自适应问题日益受到关注。然而,现有方法的迁移机制主要基于视觉或语言模态,忽略了视觉-语言联合引导的潜力。本文提出一种图像-文本多模态框架,即图像与文本肖像(ITportrait),用于实现三维肖像域自适应。ITportrait采用两阶段交替训练策略:第一阶段,我们提出三维艺术配对迁移(APT)方法实现图像引导的风格迁移。APT通过构建配对的逼真肖像获取精确的艺术姿态,从而帮助ITportrait实现高质量的三维风格迁移;第二阶段,我们在CLIP空间中提出三维图像-文本嵌入(ITE)方法。ITE利用阈值函数自适应控制CLIP空间中的图像或文本优化方向。综合实验表明,我们的ITportrait达到了最优(SOTA)性能,并能有效支持下游任务。所有源代码与预训练模型将向公众公开。