Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop at text: a character may need to follow a sketched root path, hit an end-effector target, or satisfy a multi-joint trajectory while still preserving the gait, style, and intent described by language. This exposes a control trade-off. A trajectory controller should be precise without overwriting the pretrained text-conditioned motion prior, yet existing solutions either duplicate large portions of the generator to regain per-layer control access or move much of the cost to test-time optimization. We introduce KV-Control, a compact attention-side control interface for frozen masked text-to-motion transformers. The key idea is to make geometric constraints available as memory inside self-attention rather than injecting them through a global pose token or enforcing them only at the output side. To support this interface, we co-design a part-tokenized motion substrate and controller: \textbf{PartVQ} learns anatomy-aligned part codebooks, T-Concat exposes each frame--part token as an attention-addressable site, and KV-Control injects control-conditioned key/value memories at every self-attention layer while preserving the pretrained query stream, text cross-attention, FFN, and all backbone weights. The resulting adapter adds only trainable injection parameters atop a shared trajectory encoder, yet tracks root and multi-joint constraints with sub-centimeter accuracy under the inherited refinement protocol while retaining text-conditioned motion quality. KV-Control reframes trajectory conditioning as lightweight memory retrieval, providing a small, precise, and transparent control interface for text-to-motion generation.


翻译:摘要:基于文本条件的3D人体运动模型现已能够根据提示生成合理的运动,但实际动画与具身智能体的工作流程很少止步于文本:角色可能需要沿手绘轨迹路径移动、触及末端执行器目标、或满足多关节轨迹约束,同时保持语言描述的步态、风格和意图。这暴露出一个控制权衡问题:轨迹控制器应在不覆盖预训练文本条件运动先验的前提下实现精确控制,然而现有解决方案要么复制生成器的大部分结构以重新获取每层控制权限,要么将大量计算开销转移至测试时优化。我们提出KV-Control,一种用于冻结遮罩文本条件运动Transformer的紧凑型注意力侧控制接口。其核心思想是将几何约束作为自注意力中的记忆单元引入,而非通过全局位姿令牌注入或仅在输出端施加约束。为支持该接口,我们协同设计了分部位令牌化的运动基元与控制器:PartVQ学习解剖对齐的部位码本,T-Concat将每帧-部位令牌暴露为可寻址注意力节点,KV-Control在每个自注意力层注入控制条件的键/值记忆,同时保留预训练的查询流、文本交叉注意力、前馈网络及所有骨干权重。该适配器仅在共享轨迹编码器之上增加可训练的注入参数,却在继承的细化协议下以亚厘米级精度追踪根关节和多关节约束,同时保持文本条件运动质量。KV-Control将轨迹条件化重构为轻量级记忆检索,为文生运动生成提供了小巧、精确且透明的控制接口。

0
下载
关闭预览

相关内容

【CVPR2025】MixerMDM:可学习的人体运动扩散模型组合
专知会员服务
10+阅读 · 2025年4月3日
【斯坦福博士论文】可控生成与编辑的三维神经表示,
专知会员服务
20+阅读 · 2024年12月8日
虚拟人运动控制策略学习方法的研究进展与展望
专知会员服务
19+阅读 · 2024年8月17日
【ICML2023】基于自然语言指令的受控文本生成
专知会员服务
29+阅读 · 2023年4月28日
【CVPR2023】NS3D:3D对象和关系的神经符号Grounding
专知会员服务
23+阅读 · 2023年3月26日
基于 Carsim 2016 和 Simulink的无人车运动控制联合仿真(三)
MaskFusion: 多运动目标实时识别、跟踪和重建
计算机视觉life
11+阅读 · 2019年4月20日
NLP中自动生产文摘(auto text summarization)
机器学习研究会
14+阅读 · 2017年10月10日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员