This paper reports on the GENEA Challenge 2023, in which participating teams built speech-driven gesture-generation systems using the same speech and motion dataset, followed by a joint evaluation. This year's challenge provided data on both sides of a dyadic interaction, allowing teams to generate full-body motion for an agent given its speech (text and audio) and the speech and motion of the interlocutor. We evaluated 12 submissions and 2 baselines together with held-out motion-capture data in several large-scale user studies. The studies focused on three aspects: 1) the human-likeness of the motion, 2) the appropriateness of the motion for the agent's own speech whilst controlling for the human-likeness of the motion, and 3) the appropriateness of the motion for the behaviour of the interlocutor in the interaction, using a setup that controls for both the human-likeness of the motion and the agent's own speech. We found a large span in human-likeness between challenge submissions, with a few systems rated close to human mocap. Appropriateness seems far from being solved, with most submissions performing in a narrow range slightly above chance, far behind natural motion. The effect of the interlocutor is even more subtle, with submitted systems at best performing barely above chance. Interestingly, a dyadic system being highly appropriate for agent speech does not necessarily imply high appropriateness for the interlocutor. Additional material is available via the project website at https://svito-zar.github.io/GENEAchallenge2023/ .
翻译:本文报告了GENEA 2023挑战赛的成果。参赛团队基于相同的语音和动作数据集构建了语音驱动的手势生成系统,并进行了联合评估。本年度的挑战提供了双人交互双方的完整数据,使团队能够根据智能体的语音(文本和音频)以及对话者的语音和动作,为其生成全身运动。我们联合保留动作捕捉数据,在多项大规模用户研究中评估了12个提交系统和2个基线系统。研究聚焦于三个方面:1)动作的人类相似度;2)在控制动作人类相似度的前提下,动作对智能体自身语音的适配度;3)在同时控制动作人类相似度和智能体自身语音的条件下,动作对交互中对话者行为的适配度。我们发现挑战提交系统之间的人类相似度差异显著,少数系统评分接近人体动作捕捉数据。适配度问题远未解决,多数系统表现集中在略高于随机水平的狭窄区间,远低于自然动作。对话者效应更为微妙,提交系统的最佳表现仅略高于随机水平。值得注意的是,一个双人系统若对智能体语音具有高度适配性,并不必然意味着其对对话者也具有高度适配性。补充材料可通过项目网站获取:https://svito-zar.github.io/GENEAchallenge2023/