Real-world objects perform complex motions that involve multiple independent motion components. For example, while talking, a person continuously changes their expressions, head, and body pose. In this work, we propose a novel method to decompose motion in videos by using a pretrained image GAN model. We discover disentangled motion subspaces in the latent space of widely used style-based GAN models that are semantically meaningful and control a single explainable motion component. The proposed method uses only a few $(\approx10)$ ground truth video sequences to obtain such subspaces. We extensively evaluate the disentanglement properties of motion subspaces on face and car datasets, quantitatively and qualitatively. Further, we present results for multiple downstream tasks such as motion editing, and selective motion transfer, e.g. transferring only facial expressions without training for it.
翻译:真实世界中的物体执行涉及多个独立运动分量的复杂运动。例如,在说话时,人会持续改变其表情、头部和身体姿态。在本文中,我们提出了一种新颖的方法,利用预训练的图像生成对抗网络模型对视频中的运动进行解耦。我们在广泛使用的基于风格的生成对抗网络模型的潜在空间中发现了具有语义意义且控制单一可解释运动分量的解耦运动子空间。所提出的方法仅需少量(约10个)真实视频序列即可获得此类子空间。我们在人脸和汽车数据集上对运动子空间的解耦特性进行了定量和定性评估。此外,我们还展示了多项下游任务的结果,例如运动编辑和选择性运动迁移(如无需训练即可仅迁移面部表情)。