In this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 13.7M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 96K motion sequences from massive scenes. Besides, Motion-X provides 13.7M frame-level whole-body pose descriptions and 96K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery.
翻译:本文提出Motion-X——一个大规模三维可表达全身运动数据集。现有运动数据集主要包含仅身体姿态的数据,缺乏面部表情、手部手势及细粒度姿态描述。此外,这类数据集主要采集自有限实验室场景,并依赖人工标注文本描述,极大限制了其可扩展性。为克服这些局限,我们开发了一套全身运动与文本标注流程,可自动从单视角或多视角视频中标注运动信息,并为每个视频提供全面语义标签、为每帧提供细粒度全身姿态描述。该流程具有高精度、低成本及可扩展性,适用于进一步研究。基于此,我们构建了Motion-X数据集,包含1370万组精确三维全身姿态标注(即SMPL-X),覆盖来自海量场景的9.6万条运动序列。此外,Motion-X提供1370万组帧级全身姿态描述和9.6万条序列级语义标签。综合实验表明,该标注流程具有准确性,且Motion-X在提升可表达性、多样性及自然性运动生成,以及三维全身人体网格重建方面具有显著优势。