DiffusionAvatars synthesizes a high-fidelity 3D head avatar of a person, offering intuitive control over both pose and expression. We propose a diffusion-based neural renderer that leverages generic 2D priors to produce compelling images of faces. For coarse guidance of the expression and head pose, we render a neural parametric head model (NPHM) from the target viewpoint, which acts as a proxy geometry of the person. Additionally, to enhance the modeling of intricate facial expressions, we condition DiffusionAvatars directly on the expression codes obtained from NPHM via cross-attention. Finally, to synthesize consistent surface details across different viewpoints and expressions, we rig learnable spatial features to the head's surface via TriPlane lookup in NPHM's canonical space. We train DiffusionAvatars on RGB videos and corresponding tracked NPHM meshes of a person and test the obtained avatars in both self-reenactment and animation scenarios. Our experiments demonstrate that DiffusionAvatars generates temporally consistent and visually appealing videos for novel poses and expressions of a person, outperforming existing approaches.
翻译:DiffusionAvatars 能够合成人物高保真三维头部化身,并支持对姿态和表情的直观控制。我们提出了一种基于扩散的神经渲染器,利用通用二维先验生成逼真的人脸图像。为提供表情和头部姿态的粗略引导,我们从目标视角渲染神经参数化头部模型(NPHM),该模型作为人物的代理几何结构。此外,为增强对复杂面部表情的建模,我们通过交叉注意力机制,直接利用NPHM提取的表情编码对DiffusionAvatars进行条件约束。最后,为在不同视角和表情下合成一致的表面细节,我们在NPHM的规范空间中通过TriPlane查找,将可学习的空间特征绑定至头部表面。我们基于人物的RGB视频及对应追踪的NPHM网格训练DiffusionAvatars,并在自重构与动画生成场景中测试所得化身。实验表明,DiffusionAvatars能为人物的新奇姿态与表情生成时间一致且视觉吸引人的视频,其性能优于现有方法。