We introduce Diffusion Parametric Head Models (DPHMs), a generative model that enables robust volumetric head reconstruction and tracking from monocular depth sequences. While recent volumetric head models, such as NPHMs, can now excel in representing high-fidelity head geometries, tracking and reconstructing heads from real-world single-view depth sequences remains very challenging, as the fitting to partial and noisy observations is underconstrained. To tackle these challenges, we propose a latent diffusion-based prior to regularize volumetric head reconstruction and tracking. This prior-based regularizer effectively constrains the identity and expression codes to lie on the underlying latent manifold which represents plausible head shapes. To evaluate the effectiveness of the diffusion-based prior, we collect a dataset of monocular Kinect sequences consisting of various complex facial expression motions and rapid transitions. We compare our method to state-of-the-art tracking methods and demonstrate improved head identity reconstruction as well as robust expression tracking.
翻译:我们提出扩散参数化头部模型(Diffusion Parametric Head Models, DPHMs),这是一种生成模型,能够从单目深度序列实现鲁棒的体积头部重建与跟踪。尽管近年来的体积头部模型(如NPHMs)在表示高保真头部几何结构方面已表现出色,但从真实世界单视角深度序列中跟踪和重建头部仍极具挑战性,因为对部分遮挡和噪声观测数据的拟合是欠约束的。为解决这些问题,我们提出一种基于潜在扩散的先验知识来正则化体积头部重建与跟踪过程。该基于先验的正则化器有效约束了身份编码和表情编码,使其位于表征合理头部形状的潜在流形上。为评估基于扩散先验的有效性,我们收集了一个包含各种复杂面部表情运动和快速切换的单目Kinect深度序列数据集。与最先进的跟踪方法相比,我们的方法在头部身份重建和鲁棒表情跟踪方面均表现出更优性能。