Generating dances that are both lifelike and well-aligned with music continues to be a challenging task in the cross-modal domain. This paper introduces PopDanceSet, the first dataset tailored to the preferences of young audiences, enabling the generation of aesthetically oriented dances. And it surpasses the AIST++ dataset in music genre diversity and the intricacy and depth of dance movements. Moreover, the proposed POPDG model within the iDDPM framework enhances dance diversity and, through the Space Augmentation Algorithm, strengthens spatial physical connections between human body joints, ensuring that increased diversity does not compromise generation quality. A streamlined Alignment Module is also designed to improve the temporal alignment between dance and music. Extensive experiments show that POPDG achieves SOTA results on two datasets. Furthermore, the paper also expands on current evaluation metrics. The dataset and code are available at https://github.com/Luke-Luo1/POPDG.
翻译:摘要:生成既逼真又与音乐高度契合的舞蹈,仍是跨模态领域中的一项挑战性任务。本文首次提出面向年轻观众偏好的数据集PopDanceSet,能够生成具有审美导向的舞蹈,并在音乐体裁多样性、舞蹈动作的复杂性与深度方面超越AIST++数据集。此外,在iDDPM框架下提出的POPDG模型增强了舞蹈多样性,并通过空间增强算法强化人体关节间的空间物理连接,确保多样性提升不损害生成质量。本文还设计了一个精简的对齐模块,以优化舞蹈与音乐之间的时间对齐。广泛实验表明,POPDG在两个数据集上均取得了最优结果。同时,本文对现有评估指标进行了扩展。数据集与代码已开源至https://github.com/Luke-Luo1/POPDG。