Can we make virtual characters in a scene interact with their surrounding objects through simple instructions? Is it possible to synthesize such motion plausibly with a diverse set of objects and instructions? Inspired by these questions, we present the first framework to synthesize the full-body motion of virtual human characters performing specified actions with 3D objects placed within their reach. Our system takes textual instructions specifying the objects and the associated intentions of the virtual characters as input and outputs diverse sequences of full-body motions. This contrasts existing works, where full-body action synthesis methods generally do not consider object interactions, and human-object interaction methods focus mainly on synthesizing hand or finger movements for grasping objects. We accomplish our objective by designing an intent-driven fullbody motion generator, which uses a pair of decoupled conditional variational auto-regressors to learn the motion of the body parts in an autoregressive manner. We also optimize the 6-DoF pose of the objects such that they plausibly fit within the hands of the synthesized characters. We compare our proposed method with the existing methods of motion synthesis and establish a new and stronger state-of-the-art for the task of intent-driven motion synthesis.
翻译:我们能否通过简单的指令让场景中的虚拟角色与其周围的物体进行交互?能否在多样化的物体和指令下合成出合理可信的运动?受这些问题的启发,我们提出了首个能够合成虚拟人类角色对可触及范围内的3D物体执行指定动作的全身运动框架。该系统将指定物体及其关联的虚拟角色意图的文本指令作为输入,输出多样化的全身运动序列。这与现有研究形成鲜明对比:全身动作合成方法通常不考虑物体交互,而人-物交互方法主要聚焦于抓取物体时的手部或手指运动合成。为实现这一目标,我们设计了一个意图驱动的全身运动生成器,该生成器采用一对解耦的条件变分自回归器,以自回归方式学习身体各部位的运动。同时,我们优化物体的6自由度位姿,使其能合理贴合于合成角色的手中。将所提方法与现有运动合成方法对比,我们在意图驱动运动合成任务上建立了更强的新标杆。