In this paper, we introduce a neural rendering pipeline for transferring the facial expressions, head pose, and body movements of one person in a source video to another in a target video. We apply our method to the challenging case of Sign Language videos: given a source video of a sign language user, we can faithfully transfer the performed manual (e.g., handshape, palm orientation, movement, location) and non-manual (e.g., eye gaze, facial expressions, mouth patterns, head, and body movements) signs to a target video in a photo-realistic manner. Our method can be used for Sign Language Anonymization, Sign Language Production (synthesis module), as well as for reenacting other types of full body activities (dancing, acting performance, exercising, etc.). We conduct detailed qualitative and quantitative evaluations and comparisons, which demonstrate the particularly promising and realistic results that we obtain and the advantages of our method over existing approaches.
翻译:本文提出了一种神经渲染管道,用于将源视频中一个人的面部表情、头部姿态和身体运动迁移至目标视频中的另一个人。我们将该方法应用于具有挑战性的手语视频场景:给定手语使用者的源视频,我们能够以照片级真实感的方式,忠实转移所执行的手部(如手形、手掌朝向、运动、位置)和非手部(如眼神注视、面部表情、口型模式、头部及身体运动)手语信号。该方法可用于手语匿名化、手语生成(合成模块),以及重现其他类型的全身活动(舞蹈、表演、锻炼等)。我们开展了详细的定性与定量评估及对比,结果表明该方法获得了极具前景且逼真的效果,且优于现有方法。