The ability to create realistic, animatable and relightable head avatars from casual video sequences would open up wide ranging applications in communication and entertainment. Current methods either build on explicit 3D morphable meshes (3DMM) or exploit neural implicit representations. The former are limited by fixed topology, while the latter are non-trivial to deform and inefficient to render. Furthermore, existing approaches entangle lighting in the color estimation, thus they are limited in re-rendering the avatar in new environments. In contrast, we propose PointAvatar, a deformable point-based representation that disentangles the source color into intrinsic albedo and normal-dependent shading. We demonstrate that PointAvatar bridges the gap between existing mesh- and implicit representations, combining high-quality geometry and appearance with topological flexibility, ease of deformation and rendering efficiency. We show that our method is able to generate animatable 3D avatars using monocular videos from multiple sources including hand-held smartphones, laptop webcams and internet videos, achieving state-of-the-art quality in challenging cases where previous methods fail, e.g., thin hair strands, while being significantly more efficient in training than competing methods.
翻译:从日常视频序列中创建逼真、可驱动且可重光照的头像,将为人机交互与娱乐领域带来广泛应用。现有方法要么基于显式三维可变形网格(3DMM),要么依赖神经隐式表示。前者受限于固定拓扑结构,后者难以变形且渲染效率低下。此外,现有方法在颜色估计中混淆了光照信息,因此无法在虚拟环境中重新渲染头像。为此,我们提出PointAvatar——一种可变形点云表示方法,将源颜色解耦为本征反照率与法向相关着色。我们证明,PointAvatar弥合了现有网格与隐式表示之间的鸿沟,在保持拓扑灵活性、易变形性和高效渲染的同时,实现了高质量的几何与外观建模。实验表明,本方法能够利用从手持智能手机、笔记本电脑摄像头及网络视频等多元来源获取的单目视频生成可驱动三维头像,在薄发丝等先前方法失效的挑战性场景中达到当前最优质量,且训练效率显著优于对比方法。