In this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural descriptions. We show that a simple CNN-based VQ-VAE with commonly used training recipes (EMA and Code Reset) allows us to obtain high-quality discrete representations. For GPT, we incorporate a simple corruption strategy during the training to alleviate training-testing discrepancy. Despite its simplicity, our T2M-GPT shows better performance than competitive approaches, including recent diffusion-based approaches. For example, on HumanML3D, which is currently the largest dataset, we achieve comparable performance on the consistency between text and generated motion (R-Precision), but with FID 0.116 largely outperforming MotionDiffuse of 0.630. Additionally, we conduct analyses on HumanML3D and observe that the dataset size is a limitation of our approach. Our work suggests that VQ-VAE still remains a competitive approach for human motion generation.
翻译:本研究探索了一种基于矢量量化变分自编码器(VQ-VAE)和生成式预训练Transformer(GPT)的简单且必要的条件生成框架,用于从文本描述生成人体运动。研究表明,采用常用训练策略(EMA和Code Reset)的简单CNN架构VQ-VAE可获得高质量离散表示。对于GPT,我们在训练过程中引入简单的损坏策略以缓解训练-测试差异。尽管方法简单,T2M-GPT在性能上优于包括近期扩散方法在内的竞争性方法。例如,在目前最大数据集HumanML3D上,我们在文本与生成运动的一致性(R精度)上取得可比性能,但FID指标0.116显著优于MotionDiffuse的0.630。此外,通过对HumanML3D的分析,我们观察到数据集规模限制了方法性能。本研究表明VQ-VAE仍是人体运动生成领域的有效方法。