While the recent advances in research on video reenactment have yielded promising results, the approaches fall short in capturing the fine, detailed, and expressive facial features (e.g., lip-pressing, mouth puckering, mouth gaping, and wrinkles) which are crucial in generating realistic animated face videos. To this end, we propose an end-to-end expressive face video encoding approach that facilitates data-efficient high-quality video re-synthesis by optimizing low-dimensional edits of a single Identity-latent. The approach builds on StyleGAN2 image inversion and multi-stage non-linear latent-space editing to generate videos that are nearly comparable to input videos. While existing StyleGAN latent-based editing techniques focus on simply generating plausible edits of static images, we automate the latent-space editing to capture the fine expressive facial deformations in a sequence of frames using an encoding that resides in the Style-latent-space (StyleSpace) of StyleGAN2. The encoding thus obtained could be super-imposed on a single Identity-latent to facilitate re-enactment of face videos at $1024^2$. The proposed framework economically captures face identity, head-pose, and complex expressive facial motions at fine levels, and thereby bypasses training, person modeling, dependence on landmarks/ keypoints, and low-resolution synthesis which tend to hamper most re-enactment approaches. The approach is designed with maximum data efficiency, where a single $W+$ latent and 35 parameters per frame enable high-fidelity video rendering. This pipeline can also be used for puppeteering (i.e., motion transfer).
翻译:尽管近期在视频重演研究方面取得了有希望的进展,但现有方法在捕捉精细、细节丰富且富有表现力的面部特征(例如,抿唇、噘嘴、张口及皱纹)方面仍显不足,而这些特征对于生成逼真的动画面部视频至关重要。为此,我们提出了一种端到端的表现力面部视频编码方法,通过优化基于单一身份潜在向量的低维编辑实现数据高效的高质量视频重合成。该方法建立在StyleGAN2图像反转和多阶段非线性潜在空间编辑基础上,生成与输入视频几乎可比的视频。当前基于StyleGAN潜在向量的编辑技术主要集中于简单生成静态图像的合理编辑,而我们利用位于StyleGAN2样式潜在空间(StyleSpace)中的编码,自动化潜在空间编辑以捕捉连续帧中精细的面部表现性变形。由此获得的编码可叠加于单一身份潜在向量上,支持以$1024^2$分辨率进行面部视频重演。所提出的框架以经济方式在精细层面上捕捉面部身份、头部姿态及复杂的表现性面部运动,从而绕过了训练、人物建模、对关键点/标志点的依赖以及低分辨率合成等阻碍大多数重演方法的因素。该方法以最大限度数据效率设计,其中单一的$W+$潜在向量和每帧35个参数即可实现高保真视频渲染。该流水线还可用于木偶操控(即运动迁移)。