3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent observations of human behavior. Existing approaches, however, either map high-level text descriptions to motion or rely solely on video conditioning, leaving a gap between generated dynamics and real-world motion statistics. We introduce MotionDuet, a multimodal framework that aligns motion generation with the distribution of video-derived representations. In this dual-conditioning paradigm, video cues extracted from a pretrained model (e.g., VideoMAE) ground low-level motion dynamics, while textual prompts provide semantic intent. To bridge the distribution gap across modalities, we propose Dual-stream Unified Encoding and Transformation (DUET) and a Distribution-Aware Structural Harmonization (DASH) loss. DUET fuses video-informed cues into the motion latent space via unified encoding and dynamic attention, while DASH aligns motion trajectories with both distributional and structural statistics of video features. An auto-guidance mechanism further balances textual and visual signals by leveraging a weakened copy of the model, enhancing controllability without sacrificing diversity. Extensive experiments demonstrate that MotionDuet generates realistic and controllable human motions, surpassing strong state-of-the-art baselines.
翻译:3D人体运动生成在电影、动画、游戏及具身智能领域具有关键作用。传统3D运动合成依赖昂贵的动作捕捉技术,而近期研究表明,2D视频可提供丰富且时间连贯的人体行为观测数据。然而,现有方法要么将高层级文本描述映射为运动,要么完全依赖视频条件,导致生成的运动动力学与真实世界运动统计之间存在差距。本文提出MotionDuet,一种将运动生成与视频衍生表征分布对齐的多模态框架。在该双条件范式下,从预训练模型(如VideoMAE)提取的视频线索为底层运动动力学提供约束,而文本提示则提供语义意图。为弥合模态间的分布差异,我们提出双流统一编码与变换(DUET)机制及分布感知结构协调(DASH)损失函数。DUET通过统一编码和动态注意力将视频信息线索融合进运动潜在空间,DASH则使运动轨迹与视频特征在分布统计和结构统计层面保持对齐。自动引导机制通过使用模型弱化副本进一步平衡文本与视觉信号,在保持多样性的同时提升可控性。大量实验表明,MotionDuet能够生成真实可控的人体运动,超越现有最先进基线方法。