Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leveraging implicit bridging between diffusion models, with the content preservation, however, limited to simple data patterns. We address this by imposing biased sampling in backward diffusion while maintaining the domain independence in the training stage. We construct the bias from the source domain keyframes and apply them as the gradient of content constraints, yielding a framework with keyframe manifold constraint gradients (KMCGs). Our validation demonstrates the success of training separate models to transfer between as many as ten dance motion styles. Comprehensive experiments find a significant improvement in preserving motion contents in comparison to baseline and ablative diffusion-based style transfer models. In addition, we perform a human study for a subjective assessment of the quality of generated dance motions. The results validate the competitiveness of KMCGs.
翻译:当前运动风格迁移系统的训练依赖跨风格域的一致性损失来保留内容,这阻碍了其在大规模域和私有数据上的可扩展应用。近期图像迁移研究通过利用扩散模型之间的隐式桥接展示了各域独立训练的潜力,但内容保留仅限于简单数据模式。我们通过在反向扩散过程中施加有偏采样,同时保持训练阶段的域独立性来解决这一问题。我们基于源域关键帧构建偏置,并将其作为内容约束的梯度,由此提出一种关键帧流形约束梯度(KMCGs)框架。实验验证表明,该框架成功训练了独立模型在多达十种舞蹈运动风格之间进行迁移。全面实验发现,与基线及消融的基于扩散的风格迁移模型相比,本方法在保留运动内容方面有显著提升。此外,我们通过人工研究对生成舞蹈运动的质量进行主观评估,结果验证了KMCGs的竞争力。