Head avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expressions and postures. Existing methods, categorized into 2D-based warping, mesh-based, and neural rendering approaches, present challenges in maintaining multi-view consistency, incorporating non-facial information, and generalizing to new identities. In this paper, we propose a framework named GPAvatar that reconstructs 3D head avatars from one or several images in a single forward pass. The key idea of this work is to introduce a dynamic point-based expression field driven by a point cloud to precisely and effectively capture expressions. Furthermore, we use a Multi Tri-planes Attention (MTA) fusion module in the tri-planes canonical field to leverage information from multiple input images. The proposed method achieves faithful identity reconstruction, precise expression control, and multi-view consistency, demonstrating promising results for free-viewpoint rendering and novel view synthesis.
翻译:头部化身重建是虚拟现实、在线会议、游戏及影视行业的关键应用,近年来受到计算机视觉领域的广泛关注。该领域的核心目标是忠实重建头部化身并精确控制其表情与姿态。现有方法分为基于二维变形、网格驱动和神经渲染三大类,但均存在多视角一致性不足、非面部信息融合困难以及新身份泛化能力有限等问题。本文提出名为GPAvatar的框架,通过单次前向传播即可从单张或多张图像重建三维头部化身。其核心创新在于引入由点云驱动的动态点基表情场,以精确高效地捕捉面部表情;同时,在三维平面规范场中采用多平面注意力融合模块(MTA),以充分利用多输入图像的信息。该方法实现了身份特征的忠实重建、表情的精确控制及多视角一致性,在自由视点渲染和新视角合成任务中展现出显著优势。