Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat gestures. This limits semantic grounding, speech-motion alignment, and kinematic smoothness. We propose \emph{DuoGesture}, a neuro-inspired and biomechanically informed dual-stream approach that decomposes co-speech gesture synthesis into coupled semantic and beat streams. The two streams are coordinated by a \emph{Semantic Variational Information Bottleneck}, a stochastic frame-level gate that learns when semantic gestures should override rhythmic beat motion. The semantic stream is controlled by \emph{Motion-Grounded Semantic Conditioning}, which replaces purely linguistic word embeddings with motion-language representations to provide motion-aligned semantic priors for long-tailed lexical triggers of gestures. The beat stream is further regularised by an \emph{Inertial Beat Prior}, an anthropometry-weighted arm-chain module that reduces jitter and improves rhythmic consistency without constraining semantic frames. Objective evaluations and subjective experiments show that DuoGesture outperforms strong holistic baselines, while component ablations confirm the complementary roles of semantic grounding, stochastic stream selection, and biomechanical regularisation.
翻译:共语手势生成需要同时具备语义表达性与生物力学合理的节律运动。现有整体手势模型将基于词汇的语义手势与频繁的韵律对齐拍击手势混合处理,这限制了语义基础、语音-运动对齐以及运动学平滑性。我们提出DuoGesture——一种神经启发且基于生物力学知识的双流方法,将共语手势合成分解为耦合的语义流与节律流。两流通过语义变分信息瓶颈协调——这是一种随机帧级门控机制,可学习语义手势何时应覆盖节律性拍击运动。语义流由运动基础语义条件控制,该机制用运动-语言联合表征替代纯语言词嵌入,为手势触发词的长尾词汇提供运动对齐的语义先验。节律流进一步通过惯性节拍先验进行正则化——该先验基于人体测量学的臂链模块,在减少抖动的同时提升节律一致性而不约束语义帧。客观评估与主观实验表明,DuoGesture优于强整体基线模型,而消融实验证实了语义基础、随机流选择与生物力学正则化的互补作用。