Pose transfer of human videos aims to generate a high fidelity video of a target person imitating actions of a source person. A few studies have made great progress either through image translation with deep latent features or neural rendering with explicit 3D features. However, both of them rely on large amounts of training data to generate realistic results, and the performance degrades on more accessible internet videos due to insufficient training frames. In this paper, we demonstrate that the dynamic details can be preserved even trained from short monocular videos. Overall, we propose a neural video rendering framework coupled with an image-translation-based dynamic details generation network (D2G-Net), which fully utilizes both the stability of explicit 3D features and the capacity of learning components. To be specific, a novel texture representation is presented to encode both the static and pose-varying appearance characteristics, which is then mapped to the image space and rendered as a detail-rich frame in the neural rendering stage. Moreover, we introduce a concise temporal loss in the training stage to suppress the detail flickering that is made more visible due to high-quality dynamic details generated by our method. Through extensive comparisons, we demonstrate that our neural human video renderer is capable of achieving both clearer dynamic details and more robust performance even on accessible short videos with only 2k - 4k frames.
翻译:人体视频的姿态迁移旨在生成目标人物模仿源人物动作的高保真视频。已有研究通过深度潜特征图像翻译或显式3D特征神经渲染取得了显著进展。然而,这些方法依赖大量训练数据生成逼真结果,且在训练帧数不足的互联网视频上性能下降。本文证明,即使从短单目视频训练,动态细节仍可被保留。我们提出一种耦合图像翻译动态细节生成网络(D2G-Net)的神经视频渲染框架,充分结合显式3D特征的稳定性与学习组件的表达能力。具体而言,提出一种新颖纹理表示方法,同时编码静态与姿态变化的外观特征,该表示在神经渲染阶段被映射至图像空间并生成富含细节的帧。此外,我们在训练阶段引入简洁的时间损失,以抑制因高质量动态细节生成而更显著的细节闪烁问题。通过大量对比实验证明,即使在仅含2k-4k帧的短互联网视频上,所提神经人体视频渲染器仍能实现更清晰的动态细节与更鲁棒的迁移性能。