Text-to-motion diffusion models can generate realistic animations from text prompts, but do not support fine-grained motion editing controls. In this paper, we present a method for using natural language to iteratively specify local edits to existing character animations, a task that is common in most computer animation workflows. Our key idea is to represent a space of motion edits using a set of kinematic motion editing operators (MEOs) whose effects on the source motion is well-aligned with user expectations. We provide an algorithm that leverages pre-existing language models to translate textual descriptions of motion edits into source code for programs that define and execute sequences of MEOs on a source animation. We execute MEOs by first translating them into keyframe constraints, and then use diffusion-based motion models to generate output motions that respect these constraints. Through a user study and quantitative evaluation, we demonstrate that our system can perform motion edits that respect the animator's editing intent, remain faithful to the original animation (it edits the original animation, but does not dramatically change it), and yield realistic character animation results.
翻译:文本到运动扩散模型能够根据文本提示生成逼真的动画,但不支持细粒度的运动编辑控制。本文提出一种利用自然语言迭代式指定现有角色动画局部编辑的方法,该任务在大多数计算机动画工作流程中均属常见。我们的核心思想是使用一组运动学运动编辑算子来表征运动编辑空间,这些算子对源运动产生的效果与用户预期高度一致。我们提供一种算法,利用预训练语言模型将运动编辑的文本描述转换为程序源代码,这些程序定义并在源动画上执行MEO序列。我们通过先将MEO转换为关键帧约束,再使用基于扩散的运动模型生成符合这些约束的输出运动来执行MEO。通过用户研究和定量评估,我们证明本系统能够实现符合动画师编辑意图、忠实于原始动画(仅对原动画进行编辑而非彻底改变)、并产生逼真角色动画效果的运动编辑。