Human motion transfer (HMT) aims to generate a video clip for the target subject by imitating the source subject's motion. Although previous methods have achieved remarkable results in synthesizing good-quality videos, those methods omit the effects of individualized motion information from the source and target motions, \textit{e.g.}, fine and high-frequency motion details, on the realism of the motion in the generated video. To address this problem, we propose an identity-preserved HMT network (\textit{IDPres}), which follows the pipeline of the skeleton-based method. \textit{IDpres} takes the individualized motion and skeleton information to enhance motion representations and improve the reality of motions in the generated videos. With individualized motion, our method focuses on fine-grained disentanglement and synthesis of motion. In order to improve the representation capability in latent space and facilitate the training of \textit{IDPres}, we design a training scheme, which allows \textit{IDPres} to disentangle different representations simultaneously and control them to synthesize ideal motions accurately. Furthermore, to our best knowledge, there are no available metrics for evaluating the proportion of identity information (both individualized motion and skeleton information) in the generated video. Therefore, we propose a novel quantitative metric called Identity Score (\textit{IDScore}) based on gait recognition. We also collected a dataset with 101 subjects' solo-dance videos from the public domain, named $Dancer101$, to evaluate the method. The comprehensive experiments show the proposed method outperforms state-of-the-art methods in terms of reconstruction accuracy and realistic motion.
翻译:人体动作迁移(HMT)旨在通过模仿源主体的动作,生成目标主体的视频片段。尽管先前方法在合成高质量视频方面取得了显著成果,但这些方法忽略了源动作和目标动作中个性化运动信息(例如精细且高频的动作细节)对生成视频中动作真实性的影响。为解决此问题,我们提出了一种身份保持的HMT网络(IDPres),该网络遵循基于骨架方法的流程。IDPres利用个性化运动与骨架信息增强动作表征,并提升生成视频中动作的真实性。借助个性化运动,我们的方法专注于动作的细粒度解耦与合成。为提升潜在空间的表征能力并促进IDPres的训练,我们设计了一种训练方案,使IDPres能够同时解耦不同表征,并精确控制它们以合成理想动作。此外,据我们所知,目前尚无可用指标用于评估生成视频中身份信息(包括个性化运动与骨架信息)的比例。因此,我们基于步态识别提出了一种新颖的定量指标——身份得分(IDScore)。我们还从公共领域收集了一个包含101名主体独舞视频的数据集,命名为Dancer101,以评估该方法。综合实验表明,所提方法在重建精度与动作真实性方面均优于当前最优方法。