Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has been a promising data-driven embedding approach. However, existing methods are often with inputs of interpolated intermediate frame for continuity using basic interpolation methods with keyframes, which result in a trivial local minimum during training. In this paper, we propose a novel framework to formulate latent motion manifolds with keyframe-based constraints, from which the continuous nature of intermediate token representations is considered. Particularly, our proposed framework consists of two stages for identifying a latent motion subspace, i.e., a keyframe encoding stage and an intermediate token generation stage, and a subsequent motion synthesis stage to extrapolate and compose motion data from manifolds. Through our extensive experiments conducted on both the LaFAN1 and CMU Mocap datasets, our proposed method demonstrates both superior interpolation accuracy and high visual similarity to ground truth motions.
翻译:从稀疏关键帧推导复杂的3D运动是一个极具挑战性的问题,需要保证连续性和骨架精度。动作特征通常可以从完整的关键帧序列中准确推导,因此利用Transformer的全局上下文是一种有前景的数据驱动嵌入方法。然而,现有方法通常使用基本插值方法将关键帧插值后的中间帧作为连续性输入,这会导致训练过程中陷入平凡的局部最优。本文提出一种新颖框架,通过关键帧约束构建潜在运动流形,并考虑中间标记表示的连续特性。具体而言,我们提出的框架包含两个阶段来识别潜在运动子空间(即关键帧编码阶段和中间标记生成阶段),以及后续的运动合成阶段,用于从流形中外推和合成运动数据。通过在LaFAN1和CMU Mocap数据集上进行的广泛实验,我们提出的方法在插值精度和与真实运动的高度视觉相似性方面均表现出优越性能。