The ability to animate photo-realistic head avatars reconstructed from monocular portrait video sequences represents a crucial step in bridging the gap between the virtual and real worlds. Recent advancements in head avatar techniques, including explicit 3D morphable meshes (3DMM), point clouds, and neural implicit representation have been exploited for this ongoing research. However, 3DMM-based methods are constrained by their fixed topologies, point-based approaches suffer from a heavy training burden due to the extensive quantity of points involved, and the last ones suffer from limitations in deformation flexibility and rendering efficiency. In response to these challenges, we propose MonoGaussianAvatar (Monocular Gaussian Point-based Head Avatar), a novel approach that harnesses 3D Gaussian point representation coupled with a Gaussian deformation field to learn explicit head avatars from monocular portrait videos. We define our head avatars with Gaussian points characterized by adaptable shapes, enabling flexible topology. These points exhibit movement with a Gaussian deformation field in alignment with the target pose and expression of a person, facilitating efficient deformation. Additionally, the Gaussian points have controllable shape, size, color, and opacity combined with Gaussian splatting, allowing for efficient training and rendering. Experiments demonstrate the superior performance of our method, which achieves state-of-the-art results among previous methods.
翻译:从单目肖像视频序列中重建可动画化的逼真人头虚拟形象,是弥合虚拟世界与现实世界之间差距的关键步骤。近年来,包括显式三维形变网格(3DMM)、点云和神经隐式表示在内的人头虚拟形象技术在这一持续研究中得到了应用。然而,基于3DMM的方法受限于其固定的拓扑结构,基于点的方法因涉及大量点而面临训练负担沉重的问题,而神经隐式表示方法则在形变灵活性和渲染效率上存在局限。针对这些挑战,我们提出MonoGaussianAvatar(基于单目高斯点的人头虚拟形象),这是一种新颖的方法,它利用三维高斯点表示结合高斯形变场,从单目肖像视频中学习显式人头虚拟形象。我们使用具有可调形状的高斯点来定义人头虚拟形象,从而实现灵活的拓扑结构。这些点通过高斯形变场与目标姿态及人物表情对齐,实现高效形变。此外,高斯点具有可控的形状、大小、颜色和不透明度,并结合高斯溅射技术,使训练和渲染更高效。实验表明,我们的方法性能优越,在先前方法中达到了最先进的结果。