Human motion synthesis is an important task in computer graphics and computer vision. While focusing on various conditioning signals such as text, action class, or audio to guide the generation process, most existing methods utilize skeleton-based pose representation, requiring additional skinning to produce renderable meshes. Given that human motion is a complex interplay of bones, joints, and muscles, considering solely the skeleton for generation may neglect their inherent interdependency, which can limit the variability and precision of the generated results. To address this issue, we propose a Shape-conditioned Motion Diffusion model (SMD), which enables the generation of motion sequences directly in mesh format, conditioned on a specified target mesh. In SMD, the input meshes are transformed into spectral coefficients using graph Laplacian, to efficiently represent meshes. Subsequently, we propose a Spectral-Temporal Autoencoder (STAE) to leverage cross-temporal dependencies within the spectral domain. Extensive experimental evaluations show that SMD not only produces vivid and realistic motions but also achieves competitive performance in text-to-motion and action-to-motion tasks when compared to state-of-the-art methods.
翻译:人体运动合成是计算机图形学与计算机视觉领域的重要任务。现有方法大多聚焦于利用文本、动作类别或音频等多种条件信号来引导生成过程,并采用基于骨架的姿态表征,但需额外绑定蒙皮才能生成可渲染的网格。由于人体运动是骨骼、关节与肌肉的复杂交互作用,若仅考虑骨架进行生成,可能忽略其固有的相互依赖关系,进而限制生成结果的多样性与精确性。为解决该问题,我们提出了一种形状条件运动扩散模型(Shape-conditioned Motion Diffusion Model, SMD),该模型能够直接以网格格式生成运动序列,并以指定目标网格为条件。在SMD中,输入网格通过图拉普拉斯变换转化为谱系数以实现高效表征;随后,我们提出一种谱-时间自编码器(Spectral-Temporal Autoencoder, STAE),以充分利用谱域内的跨时间依赖性。大量实验评估表明,SMD不仅能生成逼真且生动的运动,在文本到运动与动作到运动任务中,相较于现有最优方法也取得了具有竞争力的性能。