3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent observations of human behavior. Existing approaches, however, either map high-level text descriptions to motion or rely solely on video conditioning, leaving a gap between generated dynamics and real-world motion statistics. We introduce MotionDuet, a multimodal framework that aligns motion generation with the distribution of video-derived representations. In this dual-conditioning paradigm, video cues extracted from a pretrained model (e.g., VideoMAE) ground low-level motion dynamics, while textual prompts provide semantic intent. To bridge the distribution gap across modalities, we propose Dual-stream Unified Encoding and Transformation (DUET) and a Distribution-Aware Structural Harmonization (DASH) loss. DUET fuses video-informed cues into the motion latent space via unified encoding and dynamic attention, while DASH aligns motion trajectories with both distributional and structural statistics of video features. An auto-guidance mechanism further balances textual and visual signals by leveraging a weakened copy of the model, enhancing controllability without sacrificing diversity. Extensive experiments demonstrate that MotionDuet generates realistic and controllable human motions, surpassing strong state-of-the-art baselines.
翻译:三维人体运动生成在电影、动画、游戏及具身智能领域具有关键作用。传统三维运动合成依赖昂贵动作捕捉技术,而近期研究表明二维视频可提供丰富且时序一致的人体行为观测数据。然而,现有方法或仅将高层文本描述映射为运动,或仅依赖视频条件,导致生成动态与真实运动统计量之间存在差距。本文提出MotionDuet多模态框架,该框架将运动生成过程与视频导出的表征分布对齐。在此双条件范式中,从预训练模型(如VideoMAE)提取的视频线索约束低层运动动力学,而文本提示则提供语义意图。为弥合模态间的分布差异,我们提出双流统一编码变换模块(DUET)和分布感知结构协调损失(DASH)。DUET通过统一编码与动态注意力机制将视频信息线索融合至运动隐空间,DASH则使运动轨迹与视频特征的分布统计量和结构统计量对齐。自动引导机制通过利用模型的弱化副本平衡文本与视觉信号,在保持多样性的同时增强可控性。大量实验表明,MotionDuet生成的逼真可控人体运动性能超越强基线方法。