The problem of modeling an animatable 3D human head avatar under light-weight setups is of significant importance but has not been well solved. Existing 3D representations either perform well in the realism of portrait images synthesis or the accuracy of expression control, but not both. To address the problem, we introduce a novel hybrid explicit-implicit 3D representation, Facial Model Conditioned Neural Radiance Field, which integrates the expressiveness of NeRF and the prior information from the parametric template. At the core of our representation, a synthetic-renderings-based condition method is proposed to fuse the prior information from the parametric model into the implicit field without constraining its topological flexibility. Besides, based on the hybrid representation, we properly overcome the inconsistent shape issue presented in existing methods and improve the animation stability. Moreover, by adopting an overall GAN-based architecture using an image-to-image translation network, we achieve high-resolution, realistic and view-consistent synthesis of dynamic head appearance. Experiments demonstrate that our method can achieve state-of-the-art performance for 3D head avatar animation compared with previous methods.
翻译:在轻量化设置下对可动画的三维人类头部化身进行建模具有重要意义,但尚未得到良好解决。现有三维表示方法要么在肖像图像合成逼真度方面表现优异,要么在表情控制精度方面表现突出,但难以兼顾二者。为解决该问题,我们提出一种新颖的混合显式-隐式三维表示——基于人脸模型条件的神经辐射场,该方法融合了NeRF的表达能力与参数化模板的先验信息。该表示的核心是一种基于合成渲染的条件方法,能将参数化模型的先验信息融入隐式场中,同时不影响其拓扑灵活性。此外,基于混合表示,我们有效克服了现有方法中存在的形状不一致问题,并提升了动画稳定性。通过采用基于图像到图像翻译网络的整体式GAN架构,我们实现了高分辨率、逼真且视角一致的动态头部外观合成。实验表明,与现有方法相比,本方法在三维头部化身动画领域达到了最优性能。