We introduce Efficient Motion Diffusion Model (EMDM) for fast and high-quality human motion generation. Current state-of-the-art generative diffusion models have produced impressive results but struggle to achieve fast generation without sacrificing quality. On the one hand, previous works, like motion latent diffusion, conduct diffusion within a latent space for efficiency, but learning such a latent space can be a non-trivial effort. On the other hand, accelerating generation by naively increasing the sampling step size, e.g., DDIM, often leads to quality degradation as it fails to approximate the complex denoising distribution. To address these issues, we propose EMDM, which captures the complex distribution during multiple sampling steps in the diffusion model, allowing for much fewer sampling steps and significant acceleration in generation. This is achieved by a conditional denoising diffusion GAN to capture multimodal data distributions among arbitrary (and potentially larger) step sizes conditioned on control signals, enabling fewer-step motion sampling with high fidelity and diversity. To minimize undesired motion artifacts, geometric losses are imposed during network learning. As a result, EMDM achieves real-time motion generation and significantly improves the efficiency of motion diffusion models compared to existing methods while achieving high-quality motion generation. Our code will be publicly available upon publication.
翻译:本文提出高效运动扩散模型(EMDM),用于快速、高质量的人体运动生成。当前最先进的生成式扩散模型已取得令人瞩目的成果,但难以在不牺牲生成质量的前提下实现快速生成。一方面,现有方法(如运动隐空间扩散)为提升效率在隐空间中进行扩散,但学习此类隐空间本身需要付出显著代价。另一方面,通过简单增大采样步长(例如DDIM)来加速生成,往往因难以逼近复杂的去噪分布而导致质量下降。为解决这些问题,我们提出EMDM模型,该模型能够在扩散模型的多步采样过程中捕获复杂分布,从而大幅减少采样步数并显著加速生成过程。这一目标通过条件去噪扩散GAN实现,该网络能够捕获在任意(可能较大)步长条件下受控信号约束的多模态数据分布,从而实现高保真度与多样性的少步运动采样。为最小化非期望的运动伪影,在网络学习中引入了几何约束损失。实验表明,EMDM实现了实时运动生成,在保证高质量运动生成的同时,较现有方法显著提升了运动扩散模型的效率。代码将在论文发表后开源。