Recently, digital humans for interpersonal interaction in virtual environments have gained significant attention. In this paper, we introduce a novel multi-dancer synthesis task called partner dancer generation, which involves synthesizing virtual human dancers capable of performing dance with users. The task aims to control the pose diversity between the lead dancer and the partner dancer. The core of this task is to ensure the controllable diversity of the generated partner dancer while maintaining temporal coordination with the lead dancer. This scenario varies from earlier research in generating dance motions driven by music, as our emphasis is on automatically designing partner dancer postures according to pre-defined diversity, the pose of lead dancer, as well as the accompanying tunes. To achieve this objective, we propose a three-stage framework called Dance-with-You (DanY). Initially, we employ a 3D Pose Collection stage to collect a wide range of basic dance poses as references for motion generation. Then, we introduce a hyper-parameter that coordinates the similarity between dancers by masking poses to prevent the generation of sequences that are over-diverse or consistent. To avoid the rigidity of movements, we design a Dance Pre-generated stage to pre-generate these masked poses instead of filling them with zeros. After that, a Dance Motion Transfer stage is adopted with leader sequences and music, in which a multi-conditional sampling formula is rewritten to transfer the pre-generated poses into a sequence with a partner style. In practice, to address the lack of multi-person datasets, we introduce AIST-M, a new dataset for partner dancer generation, which is publicly availiable. Comprehensive evaluations on our AIST-M dataset demonstrate that the proposed DanY can synthesize satisfactory partner dancer results with controllable diversity.
翻译:近来,虚拟环境中用于人际交互的数字人受到了广泛关注。本文提出一项新型多舞者合成任务,称为舞伴生成,即合成能够与用户共舞的虚拟人舞者。该任务旨在控制领舞者与舞伴之间的姿态多样性,核心是在保持与领舞者时间协调性的同时,确保所生成舞伴的可控多样性。与先前由音乐驱动的舞蹈动作生成研究不同,我们的重点是根据预定义的多样性、领舞者的姿态以及伴奏乐曲自动设计舞伴姿态。为实现这一目标,我们提出一个名为“与你共舞”(DanY)的三阶段框架。首先,采用三维姿态采集阶段收集大量基础舞蹈姿态作为动作生成的参考;接着,引入一个超参数,通过遮蔽姿态来协调舞者之间的相似性,避免生成过度多样或一致的序列。为避免动作僵硬,设计舞蹈预生成阶段,对这些被遮蔽的姿态进行预生成,而非填充为零。随后,采用舞蹈动作迁移阶段,结合领舞序列与音乐,重写多条件采样公式,将预生成的姿态转换为具有舞伴风格的序列。实践中,为解决多人数据集的匮乏问题,我们公开了一个用于舞伴生成的新数据集AIST-M。在该数据集上的综合评估表明,所提DanY方法能够合成可控多样性的满意舞伴结果。