In this work, we advance the neural head avatar technology to the megapixel resolution while focusing on the particularly challenging task of cross-driving synthesis, i.e., when the appearance of the driving image is substantially different from the animated source image. We propose a set of new neural architectures and training methods that can leverage both medium-resolution video data and high-resolution image data to achieve the desired levels of rendered image quality and generalization to novel views and motion. We demonstrate that suggested architectures and methods produce convincing high-resolution neural avatars, outperforming the competitors in the cross-driving scenario. Lastly, we show how a trained high-resolution neural avatar model can be distilled into a lightweight student model which runs in real-time and locks the identities of neural avatars to several dozens of pre-defined source images. Real-time operation and identity lock are essential for many practical applications head avatar systems.
翻译:本文将神经头部化身技术推进至百万像素级分辨率,重点关注跨驱动合成这一极具挑战性的任务——即驱动图像外观与动画源图像存在显著差异的场景。我们提出一组新型神经网络架构与训练方法,能够同时利用中等分辨率视频数据与高分辨率图像数据,实现目标渲染质量水平,并具备对新视角与动作的泛化能力。实验表明,所提出的架构与方法可生成令人信服的高分辨率神经化身,在跨驱动场景中性能优于现有方案。最后,我们展示了如何将训练好的高分辨率神经化身模型蒸馏为轻量级学生模型,该模型可实时运行,并将神经化身身份锁定至数十个预定义源图像。实时运行与身份锁定对于头部化身系统的诸多实际应用至关重要。