Traditional methods for constructing high-quality, personalized head avatars from monocular videos demand extensive face captures and training time, posing a significant challenge for scalability. This paper introduces a novel approach to create high quality head avatar utilizing only a single or a few images per user. We learn a generative model for 3D animatable photo-realistic head avatar from a multi-view dataset of expressions from 2407 subjects, and leverage it as a prior for creating personalized avatar from few-shot images. Different from previous 3D-aware face generative models, our prior is built with a 3DMM-anchored neural radiance field backbone, which we show to be more effective for avatar creation through auto-decoding based on few-shot inputs. We also handle unstable 3DMM fitting by jointly optimizing the 3DMM fitting and camera calibration that leads to better few-shot adaptation. Our method demonstrates compelling results and outperforms existing state-of-the-art methods for few-shot avatar adaptation, paving the way for more efficient and personalized avatar creation.
翻译:传统方法通过单目视频构建高质量个性化头部化身需要大量人脸捕捉和训练时间,这给可扩展性带来了重大挑战。本文提出一种新颖方法,仅需每位用户的单张或少量图像即可创建高质量头部化身。我们基于包含2407名受试者表情的多视角数据集,学习了一个用于3D可动画化逼真头部化身的生成模型,并将其作为先验,通过少样本图像创建个性化化身。与以往的3D感知人脸生成模型不同,我们的先验基于3DMM锚定的神经辐射场骨干网络构建,我们证明该网络通过基于少样本输入的自解码方式在化身创建中更为有效。我们还通过联合优化3DMM拟合和相机标定来处理不稳定的3DMM拟合问题,从而提升少样本适应性能。我们的方法展现了令人信服的结果,在少样本化身适应方面优于现有最先进方法,为更高效、更个性化的化身创建铺平了道路。