Transforming neutral, characterless input motions to embody the distinct style of a notable character in real time is highly compelling for character animation. This paper introduces MOCHA, a novel online motion characterization framework that transfers both motion styles and body proportions from a target character to an input source motion. MOCHA begins by encoding the input motion into a motion feature that structures the body part topology and captures motion dependencies for effective characterization. Central to our framework is the Neural Context Matcher, which generates a motion feature for the target character with the most similar context to the input motion feature. The conditioned autoregressive model of the Neural Context Matcher can produce temporally coherent character features in each time frame. To generate the final characterized pose, our Characterizer network incorporates the characteristic aspects of the target motion feature into the input motion feature while preserving its context. This is achieved through a transformer model that introduces the adaptive instance normalization and context mapping-based cross-attention, effectively injecting the character feature into the source feature. We validate the performance of our framework through comparisons with prior work and an ablation study. Our framework can easily accommodate various applications, including characterization with only sparse input and real-time characterization. Additionally, we contribute a high-quality motion dataset comprising six different characters performing a range of motions, which can serve as a valuable resource for future research.
翻译:将中性、无特征的动作输入实时转化为具有特定角色风格的生动动作,对于角色动画极具吸引力。本文提出MOCHA——一种新颖的在线动作特征化框架,能将目标角色的动作风格和身体比例迁移至输入的源动作。MOCHA首先将输入动作编码为动作特征,该特征结构化身体部位拓扑并捕捉动作依赖性,以实现有效的特征化。该框架的核心是神经上下文匹配器,其为目标角色生成与输入动作特征上下文最相似的动作特征。神经上下文匹配器的条件自回归模型能在每个时间帧生成时间一致的角色特征。为生成最终的特征化姿态,我们的特征化器网络将目标动作特征的角色特性融入输入动作特征,同时保留其上下文。这通过一个结合自适应实例归一化和基于上下文映射的交叉注意力机制的Transformer模型实现,有效将角色特征注入源特征。通过与先前工作的对比及消融实验,我们验证了该框架的性能。该框架可轻松适配多种应用场景,包括仅基于稀疏输入的特征化及实时特征化。此外,我们还贡献了一个高质量动作数据集,包含六个不同角色执行的一系列动作,可作为未来研究的宝贵资源。