Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs for data acquisition, ultimately limiting the diversity and scope of human motion. In contrast, 2D human videos offer a vast and accessible source of motion data, covering a wider range of styles and activities. In this paper, we explore the use of 2D human motion extracted from videos as an alternative data source to improve text-driven 3D motion generation. Our approach introduces a novel framework that disentangles local joint motion from global movements, enabling efficient learning of local motion priors from 2D data. We first train a single-view 2D local motion generator on a large dataset of text-2D motion pairs. Then we fine-tune the generator with 3D data, transforming it into a multi-view generator that predicts view-consistent local joint motion and root dynamics. Evaluations on the well-acknowledged dataset and novel text prompts demonstrate that our method can efficiently utilize 2D data, supporting a wider range of realistic 3D human motion generation. Our code is publicly available at https://zju3dv.github.io/Motion-2-to-3/.
翻译:文本驱动的人体运动合成在电影和游戏行业的运动设计中展现出变革潜力。现有方法通常依赖三维运动捕捉数据,此类数据需要特殊采集设备,导致数据获取成本高昂,最终限制了人体运动的多样性与覆盖广度。相比之下,二维人体视频提供了海量且易获取的运动数据源,涵盖更广泛的风格与活动类型。本文探索利用视频中提取的二维人体运动作为替代数据源,以改进文本驱动的三维运动生成。我们提出一种新型框架,将局部关节运动与全局运动解耦,从而能够从二维数据中高效学习局部运动先验。首先基于大规模文本-二维运动配对数据集训练单视角二维局部运动生成器,随后通过三维数据微调该生成器,使其转化为预测视角一致的局部关节运动与根节点动力学的多视角生成器。在公认数据集与新文本提示上的评估表明,本方法能高效利用二维数据,支持更广泛场景下的逼真三维人体运动生成。相关代码已开源至 https://zju3dv.github.io/Motion-2-to-3/。