Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to separately control the object motion and adjust camera viewpoint; and (2) motion causality, ensuring that user-driven actions trigger coherent reactions from other objects rather than merely displacing pixels. Existing methods fall short on both fronts: they entangle camera and object motion into a single tracking signal and treat motion as kinematic displacement without modeling causal relationships between object motion. We introduce MoRight, a unified framework that addresses both limitations through disentangled motion modeling. Object motion is specified in a canonical static-view and transferred to an arbitrary target camera viewpoint via temporal cross-view attention, enabling disentangled camera and object control. We further decompose motion into active (user-driven) and passive (consequence) components, training the model to learn motion causality from data. At inference, users can either supply active motion and MoRight predicts consequences (forward reasoning), or specify desired passive outcomes and MoRight recovers plausible driving actions (inverse reasoning), all while freely adjusting the camera viewpoint. Experiments on three benchmarks demonstrate state-of-the-art performance in generation quality, motion controllability, and interaction awareness.


翻译:生成运动控制视频——用户指定动作驱动物理合理的场景动态,并支持自由选择视角——需要两种能力:(1)解耦运动控制,允许用户分别控制物体运动并调整相机视角;(2)运动因果性,确保用户驱动动作能够触发其他物体的连贯响应,而非仅移动像素。现有方法在这两方面均存在不足:它们将相机与物体运动纠缠为单一跟踪信号,并将运动视为运动学位移而忽视物体间的因果关系建模。我们提出MoRight,一个通过解耦运动建模同时解决上述局限的统一框架。物体运动在规范静态视角中指定,并通过时序交叉视角注意力迁移至任意目标相机视角,实现相机与物体控制的解耦。我们进一步将运动分解为主动(用户驱动)与被动(结果)组件,训练模型从数据中学习运动因果性。推理时,用户可提供主动运动并由MoRight预测结果(正向推理),或指定期望的被动效果并由MoRight恢复合理的驱动动作(反向推理),同时自由调整相机视角。在三个基准上的实验表明,该方法在生成质量、运动可控性和交互感知方面均达到最优性能。

0
下载
关闭预览

相关内容

【CVPR2025】MixerMDM:可学习的人体运动扩散模型组合
专知会员服务
10+阅读 · 2025年4月3日
虚拟人运动控制策略学习方法的研究进展与展望
专知会员服务
19+阅读 · 2024年8月17日
自动驾驶技术解读——自动驾驶汽车决策控制系统
智能交通技术
30+阅读 · 2019年7月7日
基于 Carsim 2016 和 Simulink的无人车运动控制联合仿真(三)
一文看懂如何将深度学习应用于视频动作识别
ETP:精确时序动作定位
极市平台
13+阅读 · 2018年5月25日
MoCoGAN 分解运动和内容的视频生成
CreateAMind
18+阅读 · 2017年10月21日
李克强:智能车辆运动控制研究综述
厚势
21+阅读 · 2017年10月17日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
《人工智能赋能的适应性多功能电磁战》
专知会员服务
10+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
10+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
16+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
9+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
15+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
18+阅读 · 9月22日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员